Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs

Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs
复制标题

微小而强大:设计和实现多核 SoC 的可扩展延迟容忍度

DOI:
10.1145/3470496.3527400
复制
发表时间:
2022
期刊:
ACM
影响因子:
--
通讯作者:
Martonosi, Margaret
Martonosi, Margaret
中科院分区:
--
文献类型:
--
作者:
Orenes-Vera, Marcelo;Manocha, Aninda;Balkind, Jonathan;Gao, Fei;Aragón, Juan L.;Wentzlaff, David;Martonosi, Margaret

文献摘要

参考文献

被引文献

相似文献

现代计算系统使用显著的异构性和专门化来以可管理的能力满足性能目标。然而,存储器延迟瓶颈仍然是一个问题,特别是对于稀疏神经网络和图分析应用,其中间接存储器访问(IMA)挑战存储器层次结构。许多现有技术已经提出了减少IMA延迟的硬件和软件机制,但它们无法分析真实芯片的考虑因素,特别是在用于SoC和许多核时。在本文中,我们回顾了其中的许多技术,同时考虑了多核集成和验证。我们提出了第一个延迟容错硬件的系统实现,该硬件在不需要任何存储器层次结构或处理器块修改的情况下提供了显著的加速比。这是通过内存访问并行加载引擎(Maple)实现的,该引擎以可扩展的方式通过片上网络(NOC)集成。我们的软硬件协同设计允许程序从内核异步执行长延迟内存访问,避免了流水线停滞,并实现了更大的内存并行性(MLP)。2021年4月,我们流片出了一款多核芯片,其中包括数十个Maple实例,以实现高效的数据供应。Maple演示了内核外延迟缓解硬件的完整RTL实现,以及虚拟内存支持和针对它的自动编译。该文评估了Maple与一个运行完全SMP Linux的应用程序的双核FPGA原型的集成,并展示了与基于软件的预取和解耦相比分别可获得2.35倍和2.27倍的地理加速比。与最先进的硬件相比,它提供了1.82倍和1.72倍的地理加速比,而不是预取和解耦技术。
Modern computing systems employ significant heterogeneity and specialization to meet performance targets at manageable power. However, memory latency bottlenecks remain problematic, particularly for sparse neural network and graph analytic applications where indirect memory accesses (IMAs) challenge the memory hierarchy.Decades of prior art have proposed hardware and software mechanisms to mitigate IMA latency, but they fail to analyze real-chip considerations, especially when used in SoCs and manycores. In this paper, we revisit many of these techniques while taking into account manycore integration and verification.We present the first system implementation of latency tolerance hardware that provides significant speedups without requiring any memory hierarchy or processor tile modifications. This is achieved through a Memory Access Parallel-Load Engine (MAPLE), integrated through the Network-on-Chip (NoC) in a scalable manner. Our hardware-software co-design allows programs to perform long-latency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).In April 2021 we taped out a manycore chip that includes tens of MAPLE instances for efficient data supply. MAPLE demonstrates a full RTL implementation of out-of-core latency-mitigation hardware, with virtual memory support and automated compilation targetting it. This paper evaluates MAPLE integrated with a dual-core FPGA prototype running applications with full SMP Linux, and demonstrates geomean speedups of 2.35× and 2.27× over software-based prefetching and decoupling, respectively. Compared to state-of-the-art hardware, it provides geomean speedups of 1.82× and 1.72× over prefetching and decoupling techniques.
内存延迟和细粒度并行性对 Astronautics ZS-1 性能的影响
DOI: --
发表时间: 1990
期刊: Twenty-Third Annual Hawaii International Conference on System Sciences
影响因子: --
作者:
W. Mangione;S. Abraham;E. Davidson
通讯作者: E. Davidson
DOI: 10.1145/3133901
发表时间: 2017-10-01
影响因子: 1.8
作者:
Kjolstad, Fredrik;Kamil, Shoaib;Amarasinghe, Saman
通讯作者: Amarasinghe, Saman
OUTRIDER:通过解耦链实现高效的内存延迟容忍度
DOI: 10.1145/2000064.2000079
发表时间: 2011
期刊: 2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子: --
作者:
N. Crago;Sanjay J. Patel
通讯作者: Sanjay J. Patel
DOI: 10.1145/2925426.2926254
发表时间: 2016-06
期刊: Proceedings of the 2016 International Conference on Supercomputing
影响因子: --
作者:
S. Ainsworth;Timothy M. Jones
通讯作者: S. Ainsworth;Timothy M. Jones
DOI: --
发表时间: 1992
期刊: [1992] Proceedings the 19th Annual International Symposium on Computer Architecture
影响因子: --
作者:
W. Wulf
通讯作者: W. Wulf