IMR: High-Performance Low-Cost Multi-Ring NoCs

IMR: High-Performance Low-Cost Multi-Ring NoCs
复制标题

IMR:高性能低成本多环 NoC

DOI:
10.1109/tpds.2015.2465905
复制
发表时间:
2016-06
影响因子:
5.3
通讯作者:
Chen, Y.
Chen, Y.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xu, Z.;Chen, H.;Chong, F.;Chen, Y.

文献摘要

参考文献

被引文献

相似文献

环形拓扑是片上网络(NoC)的一种常见解决方案,但经常被批评为可扩展性差。在本文中,我们提出了一种新型的多环NoC称为隔离多环(IMR),它甚至可以支持芯片多处理器(CMP)与1,024个核心。在IMR中,任何一对核心都通过至少一个隔离环连接,因此每个数据包都可以到达目的地,而无需从一个环转移到另一个环。因此,IMR不再需要昂贵的路由器作为网格,这不仅提高了网络性能,而且还减少了硬件开销。我们利用模拟进化来设计优化的IMR拓扑结构。我们将这些IMR拓扑与九个代表性的NoC(例如,传统网格、多网格、低成本网格、多虚拟信道网格(EVC)、环面环和分层环)。我们从实验中观察到,IMR显着优于其竞争对手在饱和吞吐量和延迟在所有考虑的情况下。例如,在16 × 16 CMP中,IMR将最先进网格(EVC)的饱和吞吐量平均提高了265.29%,并将SPLASH-2应用程序跟踪上的平均数据包延迟降低了71.58%,同时减少了5.08%的面积和9.76%的功耗。在32 × 32 CMP中,IMR平均将EVC的饱和吞吐量提高了191.58%,并将SPLASH-2应用程序跟踪上的数据包延迟平均降低了23.09%,同时减少了2.86%的面积和10.81%的功耗。
A ring topology is a common solution of network-on-chip (NoC) in industry, but is frequently criticized to have poor scalability. In this paper, we present a novel type of multi-ring NoC called isolated multi-ring (IMR), which can even support chip multiprocessors (CMPs) with 1,024 cores. In IMR, any pair of cores are connected via at least one isolated ring, so that each packet can reach the destination without transferring from one ring to another. Therefore, IMR no longer needs expensive routers as mesh, which not only enhances the network performance but also reduces hardware overheads. We utilize simulated evolution to design optimized IMR topologies. We compare these IMR topologies against nine representative NoCs (e.g., traditional mesh, multi mesh, low-cost mesh, Express-virtual-channels mesh (EVC), torus ring, and hierarchical ring). We observe from experiments that IMR significantly outperforms its competitors in both saturation throughput and latency across all scenarios considered. For example, in a 16 × 16 CMP, IMR improves the saturation throughput of a state-of-the-art mesh (EVC) by 265.29 percent on average, and reduces the average packet latency on SPLASH-2 application traces by 71.58 percent, while consuming 5.08 percent less area and 9.76 percent less power. In a 32 × 32 CMP, IMR averagely improves the saturation throughput of EVC by 191.58 percent, and averagely reduces the packet latency on SPLASH-2 application traces by 23.09 percent, while consuming 2.86 percent less area and 10.81 percent less power.
DOI: 10.1109/icpp.1999.797388
发表时间: 1999-09
期刊: Proceedings of the 1999 International Conference on Parallel Processing
影响因子: --
作者:
Valentin Puente;R. Beivide;J. Gregorio;J. M. Prellezo;J. Duato;C. Izu
通讯作者: Valentin Puente;R. Beivide;J. Gregorio;J. M. Prellezo;J. Duato;C. Izu
DOI: 10.1109/hpca.2010.5416635
发表时间: 2010-04
期刊: HPCA - 16 2010 The Sixteenth International Symposium on High-Performance Computer Architecture
影响因子: --
作者:
Jason E. Miller;H. Kasture;George Kurian;Charles Gruenwald;Nathan Beckmann;Christopher Celio;J. Eastep-J.-Eas
通讯作者: Jason E. Miller;H. Kasture;George Kurian;Charles Gruenwald;Nathan Beckmann;Christopher Celio;J. Eastep-J.-Eas
DOI: 10.1109/hpca.2010.5416639
发表时间: 2010-04
期刊: HPCA - 16 2010 The Sixteenth International Symposium on High-Performance Computer Architecture
影响因子: --
作者:
Aniruddha N. Udipi;N. Muralimanohar;R. Balasubramonian
通讯作者: Aniruddha N. Udipi;N. Muralimanohar;R. Balasubramonian
DOI: 10.1145/1669112.1669145
发表时间: 2009-12
期刊: 2009 42nd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子: --
作者:
John Kim
通讯作者: John Kim
DOI: 10.1109/tc.2013.2295523
发表时间: 2015-03
影响因子: 3.7
作者:
Sheng Ma;Zhiying Wang;Zonglin Liu;Natalie D. Enright Jerger
通讯作者: Sheng Ma;Zhiying Wang;Zonglin Liu;Natalie D. Enright Jerger