The StageNet fabric for constructing resilient multicore systems

The StageNet fabric for constructing resilient multicore systems
复制标题

用于构建弹性多核系统的 StageNet 结构

DOI:
--
复制
发表时间:
2008
期刊:
2008 41st IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
S. Mahlke
S. Mahlke
中科院分区:
--
文献类型:
--
作者:
S. Gupta;Shuguang Feng;Amin Ansari;Jason A. Blome;S. Mahlke

文献摘要

被引文献

相似文献

CMOS特征尺寸的缩放长期以来一直是显著性能增益的来源。然而,电压水平的降低无法匹配这种缩放速率,导致工作温度和电流密度增加。考虑到困扰半导体器件的大多数磨损机制高度依赖于这些参数,预计未来几代技术的故障率将显著提高。因此,高可靠性和容错性,这一直是高端服务器市场感兴趣的主题,现在正在主流桌面和嵌入式系统领域得到重视。对此,流行的解决方案是使用粗粒度的冗余,例如双重/三重模块冗余。在这项工作中,我们挑战的做法,粗粒度冗余识别其无法扩展到高故障率的情况下,并调查细粒度配置的优势。为此,本文提出并评估了一个高度可重构的多核架构,名为StageNet(SN),这是设计与可靠性作为其第一类设计标准。SN依赖于复制处理器流水线阶段的可重配置网络,以最大限度地提高芯片的使用寿命,并在寿命结束时适度降低性能。我们的研究结果表明,所提出的SN架构可以执行近50%的累积工作相比,传统的多核。
Scaling of CMOS feature size has long been a source of dramatic performance gains. However, the reduction in voltage levels has not been able to match this rate of scaling, leading to increasing operating temperatures and current densities. Given that most wearout mechanisms that plague semiconductor devices are highly dependent on these parameters, significantly higher failure rates are projected for future technology generations. Consequently, high reliability and fault tolerance, which have traditionally been subjects of interest for high-end server markets, are now getting emphasis in the mainstream desktop and embedded systems space. The popular solution for this has been the use of redundancy at a coarse granularity, such as dual/triple modular redundancy. In this work, we challenge the practice of coarse-granularity redundancy by identifying its inability to scale to high failure rate scenarios and investigating the advantages of finer-grained configurations. To this end, this paper presents and evaluates a highly reconfigurable multicore architecture, named StageNet (SN), that is designed with reliability as its first class design criteria. SN relies on a reconfigurable network of replicated processor pipeline stages to maximize the useful lifetime of a chip, gracefully degrading performance towards the end of life. Our results show that the proposed SN architecture can perform nearly 50% more cumulative work compared to a traditional multicore.