Express Link Placement for NoC-Based Many-Core Platforms

Express Link Placement for NoC-Based Many-Core Platforms
复制标题

DOI:
10.1145/3337821.3337877
复制
发表时间:
2019-08
期刊:
Proceedings of the 48th International Conference on Parallel Processing
影响因子:
--
通讯作者:
Yunfan Li;Di Zhu;Lizhong Chen
Yunfan Li;Di Zhu;Lizhong Chen
中科院分区:
其他
文献类型:
--
作者:
Yunfan Li;Di Zhu;Lizhong Chen

文献摘要

相似文献

随着近年来通用处理器集成了多达数百个核,并可用于并行处理系统,设计可扩展的低延迟片上网络(NoC)以支持各种片上通信变得至关重要。减少片上延迟和提高网络可扩展性的一种有效方法是在非相邻路由器对之间添加快速链路。然而,由于芯片上的有限的总二等分带宽,增加快速链路的数量可能导致每个链路的带宽更小,从而导致网络中的分组的更高的串行化延迟。不同于以往的作品,特定于应用程序的设计或固定放置的快速链接,本文的目的是找到有效的放置的快速链接的通用处理器考虑所有可能的位置选项。我们制定的数学问题,并提出了一个有效的算法,利用初始解生成启发式和增强的候选生成器模拟退火。使用多线程PARSEC基准测试和各种合成流量模式对4x4、8x8和16x16网络进行的评估显示,与以前的工作相比,平均数据包延迟显著降低。
With the integration of up to hundreds of cores in recent general-purpose processors that can be used in parallel processing systems, it is critical to design scalable and low-latency networks-on-chip (NoCs) to support various on-chip communications. An effective way to reduce on-chip latency and improve network scalability is to add express links between pairs of non-adjacent routers. However, increasing the number of express links may result in smaller bandwidth per link due to the limited total bisection bandwidth on chip, thus leading to higher serialization latency of packets in the network. Unlike previous works on application-specific designs or on fixed placement of express links, this paper aims at finding effective placement of express links for general-purpose processors considering all the possible placement options. We formulate the problem mathematically and propose an efficient algorithm that utilizes an initial solution generation heuristic and enhanced candidate generator in simulated annealing. Evaluation on 4x4, 8x8 and 16x16 networks using multi-threaded PARSEC benchmarks and various synthetic traffic patterns shows significant reduction of average packet latency over previous works.