Virtual-Link: A Scalable Multi-Producer Multi-Consumer Message Queue Architecture for Cross-Core Communication

Virtual-Link: A Scalable Multi-Producer Multi-Consumer Message Queue Architecture for Cross-Core Communication
复制标题

DOI:
10.1109/ipdps49936.2021.00027
复制
发表时间:
2020-12
期刊:
2021 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Qinzhe Wu;J. Beard;Ashen Ekanayake;A. Gerstlauer;L. John
Qinzhe Wu;J. Beard;Ashen Ekanayake;A. Gerstlauer;L. John
中科院分区:
其他
文献类型:
--
作者:
Qinzhe Wu;J. Beard;Ashen Ekanayake;A. Gerstlauer;L. John

文献摘要

相似文献

随着每个Systemon-Chip的加工元素数量增加,跨核通信越来越成为一种瓶颈。跨核通信的典型硬件解决方案通常不灵活。尽管软件解决方案是灵活的,但它们具有性能缩放限制。正如我们将显示的那样,一个关键问题是基于软件的消息队列机制中共享状态的关键问题。本文提出了虚拟链接(VL),这是一种具有硬件支持的新型轻型通信机制,可促进M:N锁定数据移动。 VL将相干共享状态的数量减少到零。 VL通过将数据保存在快速路径(即在OnChip InterConnect中)来提供进一步的延迟益处。 VL可以在相干总线上的PES之间进行定向缓存(藏匿),从而减少了Coreto-Core通信的延迟。 VL对于流媒体数据的细粒度任务特别有效。在具有7个基准测试的完整系统模拟器上进行评估表明,VL在基于最先进的软件的通信机制上实现了$ 2.09 \ times $加速,同时将内存流量降低了61%。
Cross-core communication is increasingly a bottleneck as the number of processing elements increase per systemon-chip. Typical hardware solutions to cross-core communication are often inflexible; while software solutions are flexible, they have performance scaling limitations. A key problem, as we will show, is that of shared state in software-based message queue mechanisms. This paper proposes Virtual-Link (VL), a novel light-weight communication mechanism with hardware support to facilitate M:N lock-free data movement. VL reduces the amount of coherent shared state, which is a bottleneck for many approaches, to zero. VL provides further latency benefit by keeping data on the fast path (i.e., within the onchip interconnect). VL enables directed cache-injection (stashing) between PEs on the coherence bus, reducing the latency for coreto-core communication. VL is particularly effective for fine-grain tasks on streaming data. Evaluation on a full system simulator with 7 benchmarks shows that VL achieves a $2.09\times$ speedup over state-of-the-art software-based communication mechanisms, while reducing memory traffic by 61%.