Necromancer: enhancing system throughput by animating dead cores

Necromancer: enhancing system throughput by animating dead cores
复制标题

Necromancer:通过动画死核来提高系统吞吐量

DOI:
--
复制
发表时间:
2010
期刊:
International Symposium on Computer Architecture
影响因子:
--
通讯作者:
S. Mahlke
S. Mahlke
中科院分区:
--
文献类型:
--
作者:
Amin Ansari;Shuguang Feng;S. Gupta;S. Mahlke

文献摘要

被引文献

相似文献

在过去的几年里,积极的技术扩展到纳米范围导致了大量的可靠性挑战。与可以使用传统方案有效保护的片上缓存不同,一般的核心区域不太均匀和结构化,使得容忍缺陷成为更具挑战性的问题。由于缺乏有效的解决方案,禁用非功能核是工业上提高制造良率的常见做法,这导致系统吞吐量的显著降低。虽然一个错误的核心不能被信任正确执行程序,我们在这项工作中观察到,对于大多数缺陷,从一个有效的架构状态开始时,执行痕迹上的一个有缺陷的核心实际上粗略地类似于那些无故障的执行。鉴于这一见解,我们提出了一个强大的和异构的核心耦合执行计划,死灵法师,利用一个功能死的核心,以提高系统吞吐量,提供有关高层次的程序行为的提示。我们将传统CMP系统中的核心划分为多个组,其中每个组共享一个轻量级核心,该轻量级核心可以使用来自潜在死核心的这些执行提示来大幅加速。为了防止这个不死的核心偏离正确的执行路径太远,我们用轻量级核心动态地将体系结构状态重新配置。对于一个4核CMP系统,平均而言,我们的方法使耦合的核心实现78.5%的性能的一个完全功能的核心。这种缺陷容限和吞吐量增强分别以5.3%和8.5%的适度面积和功率开销实现。
Aggressive technology scaling into the nanometer regime has led to a host of reliability challenges in the last several years. Unlike on-chip caches, which can be efficiently protected using conventional schemes, the general core area is less homogeneous and structured, making tolerating defects a much more challenging problem. Due to the lack of effective solutions, disabling non-functional cores is a common practice in industry to enhance manufacturing yield, which results in a significant reduction in system throughput. Although a faulty core cannot be trusted to correctly execute programs, we observe in this work that for most defects, when starting from a valid architectural state, execution traces on a defective core actually coarsely resemble those of fault-free executions. In light of this insight, we propose a robust and heterogeneous core coupling execution scheme, Necromancer, that exploits a functionally dead core to improve system throughput by supplying hints regarding high-level program behavior. We partition the cores in a conventional CMP system into multiple groups in which each group shares a lightweight core that can be substantially accelerated using these execution hints from a potentially dead core. To prevent this undead core from wandering too far from the correct path of execution, we dynamically resynchronize architectural state with the lightweight core. For a 4-core CMP system, on average, our approach enables the coupled core to achieve 78.5% of the performance of a fully functioning core. This defect tolerance and throughput enhancement comes at modest area and power overheads of 5.3% and 8.5%, respectively.