课题基金 / 基金详情

CAREER: Adaptively Boosting Resilience Efficiency in the Face of Frequent, Clustered, and Diverse Faults

CAREER: Adaptively Boosting Resilience Efficiency in the Face of Frequent, Clustered, and Diverse Faults
职业:面对频繁、聚集和多样化的故障,自适应地提高弹性效率
批准号:
1253733
负责人:
Chengmo Yang
金额:
$44.95万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2013
资助国家:
美国
项目状态:
已结题
起止时间:
2013-06-01 至 2019-05-31

项目摘要

项目成果

Chengmo Yang的其他基金

相似基金

相关文献

中文摘要
翻译
虽然技术进步使研究人员能够生产性能更高、功耗更低的芯片,但我们提供这种计算能力的能力受到了硅设备对故障的日益敏感的挑战。预计在未来的计算机系统中,故障将以连续的方式发生,从硬件到应用程序的所有级别都会发生故障。此外,预计过错行为将更加多样化和不可预测。关键的问题不仅是永久性和暂时性故障,而且是在纳秒到秒的时间尺度上频繁和不规则地发生的间歇性故障。这些预测的高故障率和多样化的故障行为要求故障恢复方法的转变。当故障连续发生时,故障检测和恢复都必须以更细粒度的方式执行,恢复变得与检测一样关键。此外,由于故障持续时间差异很大,需要具有成本效益的解决方案,能够统一检测所有类型的故障,识别故障类型,然后自适应地恢复执行。为了应对这些可靠性挑战,提出的项目将在系统中融入细粒度的自适应性,并将静态提取的应用程序信息与运行时优化相结合,以指导适配决策。提出的研究包括:(1)自适应检测和检查点,能够调整检测和检查点的粒度以匹配系统可靠性水平;(2)自适应恢复,能够以最小化再次故障发生的方式执行重新执行;(3)自适应资源管理,能够监控应用和硬件可靠性水平并快速调整调度决策。
英文摘要
While technology advances allow researchers to produce chips with higher performance and lower power consumption, our ability to deliver such computational power is challenged by the increasing susceptibility of silicon devices to faults. It is expected that in future computer systems, faults will occur in a continuous manner, across all levels from hardware to application. The fault behavior is furthermore expected to be more diverse and unpredictable. Of critical concern will be not only permanent and transient faults, but also intermittent faults that occur frequently and irregularly over nanosecond to second time scales.These predicted high fault rates and diverse fault behaviors mandate a transformation in fault resilience approaches. When faults occur in a continuous manner, both fault detection and recovery must be performed in a much finer-grained manner, and recovery becomes as critical as detection. Moreover, since fault duration varies significantly, cost-effective solutions capable of uniformly detecting all types of faults, identifying the fault type, and then adaptively recovering the execution are necessary.To address these reliability challenges, the proposed project will incorporate fine-grained adaptivity into the system, and couple statically extracted application information with runtime optimizations to guide adaptation decisions. The proposed research includes: (1) adaptive detection and checkpointing, capable of adjusting detection and checkpointing granularity to match system reliability levels; (2) adaptive recovery, capable of performing re-execution in a way that minimizes the chance of another fault occurring; and (3) adaptive resource management, capable of monitoring application and hardware reliability levels and quickly adapting scheduling decisions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF: Small: Collaborative Research: Retraining-free Concurrent Test and Diagnosis in Emerging Neural Network Accelerators
  • 批准号:
    1909854
  • 项目类别:
    Standard Grant
  • 资助金额:
    $26.5万
  • 财政年份:
    2019
  • 负责人:
    Chengmo Yang
  • 依托单位:
CPS: Medium: Collaborative Research: Constantly on the Lookout: Low-Cost Sensor Enabled Explosive Detection to Protect High Density Environments
  • 批准号:
    1739390
  • 项目类别:
    Standard Grant
  • 资助金额:
    $18.0万
  • 财政年份:
    2017
  • 负责人:
    Chengmo Yang
  • 依托单位:
SHF: Small: Collaborative Research: Multi-level Non-volatile FPGA Synthesis to Empower Efficient Self-adaptive System Implementations
  • 批准号:
    1527464
  • 项目类别:
    Standard Grant
  • 资助金额:
    $25.0万
  • 财政年份:
    2015
  • 负责人:
    Chengmo Yang
  • 依托单位:
海外基金