课题基金 / 基金详情

CSR-PSCE,SM: A Holistic Design Approach to Reliability Using 3D Stacked

CSR-PSCE,SM: A Holistic Design Approach to Reliability Using 3D Stacked
CSR-PSCE,SM:使用 3D 堆叠的可靠性整体设计方法
批准号:
0834798
负责人:
Murali Annavaram
金额:
$40.29万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2008
资助国家:
美国
项目状态:
已结题
起止时间:
2008-09-01 至 2013-08-31

项目摘要

项目成果

Murali Annavaram的其他基金

相似基金

相关文献

中文摘要
翻译
信息技术产业的未来取决于设计能够容忍设备特性变化引起的错误的计算机系统。传统上,系统可靠性是通过复制关键系统组件来实现的。由于可变性引起的错误随着时间的推移缓慢发生,因此对于低成本计算平台来说,仅为提供可靠性而进行复制的成本高得令人望而却步。本研究探讨使用3D堆叠技术在3D堆叠芯片上实现冗余元件和变异性监控电路。通过使用3D堆叠,可以使用变化弹性处理技术来构建冗余计算块,该技术可能比用于构建主处理器的处理技术慢。这项研究从创新的微体系结构解决方案到利用应用程序固有的容错能力,采取了一种整体的方法来设计3D堆叠监控。在微体系结构方面,这项研究探索了无缝重新配置监控层以三种模式工作的可能性:性能辅助,当变异性引起的错误很少时,或者作为保护处理器,当变异性引起的错误开始出现时,或者作为备份处理器,当器件老化可能导致主处理基板上的不可修复的错误时。在体系结构方面,这项研究探索了一个新的异常类,称为可靠性感知异常,它允许微体系结构块引发异常,以响应可变性导致的错误。然后,这些软件可见异常可由应用程序类利用,这些应用程序类本质上是容错的,并且可以定制异常处理机制。
英文摘要
The future of information technology industry depends on designing computer systems that are tolerant of errors caused by variations in device characteristics. Traditionally system reliability is achieved by replicating critical system components. Since variability induced errors occur slowly over time, replication for the sole purpose of providing reliability is prohibitively expensive for low cost computing platforms. This research explores using 3D stacking to implement redundant components and variability monitoring circuitry on a 3D stacked die. Using 3D stacking the redundant computation blocks can be built using a variation resilient process technology that may be slower than the process technology used for building the primary processor. This research takes a holistic approach to designing the 3D stacked monitoring spanning from innovative microarchitecture solutions to exploiting application's inherent error tolerance. On the microarchitecture front, this research explores the potential for seamlessly reconfiguring the monitoring layer to act in three modes: performance assists, when variability induced errors are rare, or as guard processors, when variability induced errors begin to appear, or as backup processors, when device aging may result in irreparable errors on the primary processing substrate. On the architecture front, this research explores a new exception class called Reliability Aware Exceptions that allow microarchitecture blocks to raise an exception in response to a variability induced error. These software visible exceptions can then be exploited by application classes that are inherently error tolerant and can customized exception handling mechanisms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF: Small: ML Accelerator Cohort Architecture
  • 批准号:
    2224319
  • 项目类别:
    Standard Grant
  • 资助金额:
    $60.0万
  • 财政年份:
    2022
  • 负责人:
    Murali Annavaram
  • 依托单位:
Student Travel Support for the 2018 International Symposium on Computer Architecture (ISCA)
  • 批准号:
    1812942
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.5万
  • 财政年份:
    2018
  • 负责人:
    Murali Annavaram
  • 依托单位:
SHF:Small: Accelerating Graph Analytics Through Coordinated Storage, Memory and Computing Advances
  • 批准号:
    1719074
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2017
  • 负责人:
    Murali Annavaram
  • 依托单位:
SHF:Small: Benchmarking of Transient and Intermittent Errors and Their Application to Microarchitecture
  • 批准号:
    1219186
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2012
  • 负责人:
    Murali Annavaram
  • 依托单位:
海外基金