课题基金 / 基金详情

Collaborative Research: CSR-SMA+AES: Pro-Active Runtime Health Enhancement of Large-Scale Parallel Systems Using PROGNOSIS

Collaborative Research: CSR-SMA+AES: Pro-Active Runtime Health Enhancement of Large-Scale Parallel Systems Using PROGNOSIS
合作研究:CSR-SMA AES:使用 PROGNOSIS 主动增强大规模并行系统的运行时健康状况
批准号:
0614976
负责人:
Yanyong Zhang
金额:
$23.8万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-08-15 至 2010-07-31

项目摘要

项目成果

Yanyong Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
大规模并行系统对于应对极其重要的高要求应用程序所带来的挑战至关重要。突破硬件和软件技术的极限以获取最高性能可能会增加它们对故障的敏感度。这是由于不断增加的硬件瞬时错误、硬件设备故障和软件复杂性造成的。这些故障可能会对系统性能产生重大影响,并增加维护/操作成本,从而危及部署这些大规模系统的动机。该项目不是将故障视为异常并采取主动补救措施,而是在运行时预测故障的发生并采取主动措施来隐藏其影响。本研究有望对运行时容错基础设施的开发做出三个广泛的贡献。第一组贡献是收集和分析实际Bluegene/L系统在较长一段时间内的系统事件。第二组贡献是用于在线分析和预测不断变化的故障数据的模型,第三组贡献是关于故障感知并行作业调度和检查点。在教育方面,除了加强研究生课程和研究外,该项目还打算让本科生和妇女参与进来。在该项目中开发的工具和相关成果将在公共领域提供,并在主要期刊/会议上发表。此外,PI还将推动这些工具整合到实际系统中,以增强其容错能力。
英文摘要
Large scale parallel systems are critical to take on the challenges imposed by highly demanding applications of critical importance. Pushing the limits of hardware and software technologies to extract the maximum performance can increase their susceptibility to failures. This arises as a consequence of growing hardware transient errors, hardware device failures, and software complexity. These failures can have substantial consequences on system performance, and add to the costs of maintenance/operation, thereby putting at risk the very motivation behind deploying these large scale systems. Rather than treat failures as an exception and takereactive remedies, this project intends to anticipate their occurrence and take pro-active runtime measures to hide their impact.This research is expected to make three broad contributions towardsdeveloping a runtime fault-tolerance infrastructure.The first set of contributions is on collecting and analyzingsystem events from an actual BlueGene/L system over anextended period of time. The second set of contributions are models foronline analysis and prediction of evolving failure data.The third set of contributions are on failure-aware parallel job scheduling and checkpointing. On the educational front, in addition to enhancing graduate curriculum and research, this project intends to involve undergraduate students and women. The tools developed in this project and the related results will be made available in public domain and published in leading journals/conferences. In addition, the PIs will also push these tools to be incorporated on actual systems, to enhance their fault-toleranceabilities.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
NeTS: Small: Transmit Only: Green Communication for Dense Wireless Systems
  • 批准号:
    1423020
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.86万
  • 财政年份:
    2014
  • 负责人:
    Yanyong Zhang
  • 依托单位:
CT - ISG: ROME: Robust Measurement in Sensor Networks
  • 批准号:
    0831186
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2008
  • 负责人:
    Yanyong Zhang
  • 依托单位:
CAREER: PROSE: Providing Robustness in Systems of Embedded Sensors
  • 批准号:
    0546072
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $44.97万
  • 财政年份:
    2006
  • 负责人:
    Yanyong Zhang
  • 依托单位:
Collaborative Research: CSR---SMA+AES: PROGNOSIS to Enhance the Runtime Health of Large Scale Parallel Systems
  • 批准号:
    0509164
  • 项目类别:
    Standard Grant
  • 资助金额:
    $8.0万
  • 财政年份:
    2005
  • 负责人:
    Yanyong Zhang
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)