A Lightweight Fault-Tolerant Mechanism for Network-on-Chip

A Lightweight Fault-Tolerant Mechanism for Network-on-Chip
复制标题

DOI:
10.1109/nocs.2008.34
复制
发表时间:
2007-12
期刊:
Second ACM/IEEE International Symposium on Networks-on-Chip (nocs 2008)
影响因子:
--
通讯作者:
M. Koibuchi;Hiroki Matsutani;H. Amano;T. Pinkston
M. Koibuchi;Hiroki Matsutani;H. Amano;T. Pinkston
中科院分区:
其他
文献类型:
--
作者:
M. Koibuchi;Hiroki Matsutani;H. Amano;T. Pinkston

文献摘要

被引文献

相似文献

生存能力已成为设计使用片上数据包网络或芯片(NOCS)网络构建的多层处理器的关键因素。在本文中,我们提出了一种基于默认备份路径(DBP)的NOC的轻质故障机制(DBP),旨在在存在故障的情况下维护两个非故障路由器的网络连接以及可以连接的健康处理器核心路由器故障。该机制提供了默认路径作为某些路由器端口之间的备份,这些端口是替代数据的替代数据,以规避故障路由器内的失败组件。除了普通网络通道的最小子集外,默认备份路径内置到故障路由器形式的内部 - 在最坏的情况下 - 单向环拓扑,可为所有处理器内核提供网络范围的连接。事实证明,使用DBP机制的路由被证明是无僵硬的,只有两个虚拟通道,即使在故障场景中,常规网络将降低到不规则(任意)拓扑的情况下。评估结果表明,对于2-D网格虫洞NOC,仅需要12.6%的额外硬件资源来实施拟议的DBP机制,以便在无需芯片范围的情况下提供优雅的性能降级,而随着故障的增加数量,形成环。
Survival capability is becoming a crucial factor in designing multicore processors built with on-chip packet networks, or networks on chip (NoCs). In this paper, we propose a lightweight fault-tolerant mechanism for NoCs based on default backup paths (DBPs) designed to maintain, in the presence of failures, network connectivity of both non-faulty routers as well as healthy processor cores which may be connected to faulty routers. The mechanism provides default paths as backup between certain router ports which serve as alternative datapaths to circumvent failed components within a faulty router. Along with a minimal subset of normal network channels, the set of default backup paths internal to faulty routers form - in the worst case - a unidirectional ring topology that provides network-wide connectivity to all processor cores. Routing using the DBP mechanism is proved to be deadlock-free with only two virtual channels even for fault scenarios in which regular networks degrade to irregular (arbitrary) topologies. Evaluation results show that, for a 2-D mesh wormhole NoC, only 12.6% additional hardware resources are needed to implement the proposed DBP mechanism in order to provide graceful performance degradation without chip-wide failure as the number of faults increases to the maximum needed to form ring.