High-performance, Energy-efficient, Fault-tolerant Network-on-Chip Design Using Reinforcement Learnin

High-performance, Energy-efficient, Fault-tolerant Network-on-Chip Design Using Reinforcement Learnin
复制标题

DOI:
10.23919/date.2019.8714869
复制
发表时间:
2019-03
期刊:
2019 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
Ke Wang;A. Louri;Avinash Karanth;Razvan C. Bunescu
Ke Wang;A. Louri;Avinash Karanth;Razvan C. Bunescu
中科院分区:
其他
文献类型:
--
作者:
Ke Wang;A. Louri;Avinash Karanth;Razvan C. Bunescu

文献摘要

被引文献

相似文献

片上网络 (NoC) 正在成为多核和片上系统 (SoC) 架构的标准通信结构。随着技术不断扩展,芯片上的晶体管和线路越来越容易受到各种故障机制的影响,尤其是时序错误,从而导致片上网络的能源效率和性能恶化。处理定时错误的典型技术本质上是反应性的,在故障发生后对其做出响应。它们依赖于错误检测/纠正技术,但由于错误检测/纠正硬件不断启用,导致功耗过高和性能下降。另一方面,不加区别地禁用错误处理硬件可能会导致更多错误和侵入性重传流量。因此,挑战在于平衡错误率、数据包重传、性能和能量之间的权衡。在本文中,我们提出了一种主动容错机制,通过强化学习(RL)来优化能源效率和性能。首先,我们提出了一种新的主​​动错误处理技术,包括用于启用每个路由器错误检测/纠正硬件的动态方案和有效的重传机制。其次,我们建议使用强化学习来训练动态控制策略,与传统技术相比,其目标是提高容错能力、降低功耗和提高性能。我们的评估表明,与反应式纠错技术相比,端到端数据包延迟平均降低了 55%,能源效率提高了 64%,故障导致的重传减少了 48%。
Network-on-Chips (NoCs) are becoming the standard communication fabric for multi-core and system on a chip (SoC) architectures. As technology continues to scale, transistors and wires on the chip are becoming increasingly vulnerable to various fault mechanisms, especially timing errors, resulting in exacerbation of energy efficiency and performance for NoCs. Typical techniques for handling timing errors are reactive in nature, responding to the faults after their occurrence. They rely on error detection/correction techniques which have resulted in excessive power consumption and degraded performance, since the error detection/correction hardware is constantly enabled. On the other hand, indiscriminately disabling error handling hardware can induce more errors and intrusive retransmission traffic. Therefore, the challenge is to balance the trade-offs among error rate, packet retransmission, performance, and energy. In this paper, we propose a proactive fault-tolerant mechanism to optimize energy efficiency and performance with reinforcement learning (RL). First, we propose a new proactive error handling technique comprised of a dynamic scheme for enabling per-router error detection/correction hardware and an effective retransmission mechanism. Second, we propose the use of RL to train the dynamic control policy with the goals of providing increased fault-tolerance, reduced power consumption and improved performance as compared to conventional techniques. Our evaluation indicates that, on average, end-to-end packet latency is lowered by 55%, energy efficiency is improved by 64%, and retransmission caused by faults is reduced by 48% over the reactive error correction techniques.