Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs
Tiny but mighty: designing and realizing scalable latency tolerance for manycore SoCs
复制标题
微小而强大:设计和实现多核 SoC 的可扩展延迟容忍度
DOI:
10.1145/3470496.3527400
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Martonosi, Margaret
中科院分区:
文献类型:
--
作者:
Orenes-Vera, Marcelo;Manocha, Aninda;Balkind, Jonathan;Gao, Fei;Aragón, Juan L.;Wentzlaff, David;Martonosi, Margaret
Modern computing systems employ significant heterogeneity and specialization to meet performance targets at manageable power. However, memory latency bottlenecks remain problematic, particularly for sparse neural network and graph analytic applications where indirect memory accesses (IMAs) challenge the memory hierarchy.Decades of prior art have proposed hardware and software mechanisms to mitigate IMA latency, but they fail to analyze real-chip considerations, especially when used in SoCs and manycores. In this paper, we revisit many of these techniques while taking into account manycore integration and verification.We present the first system implementation of latency tolerance hardware that provides significant speedups without requiring any memory hierarchy or processor tile modifications. This is achieved through a Memory Access Parallel-Load Engine (MAPLE), integrated through the Network-on-Chip (NoC) in a scalable manner. Our hardware-software co-design allows programs to perform long-latency memory accesses asynchronously from the core, avoiding pipeline stalls, and enabling greater memory parallelism (MLP).In April 2021 we taped out a manycore chip that includes tens of MAPLE instances for efficient data supply. MAPLE demonstrates a full RTL implementation of out-of-core latency-mitigation hardware, with virtual memory support and automated compilation targetting it. This paper evaluates MAPLE integrated with a dual-core FPGA prototype running applications with full SMP Linux, and demonstrates geomean speedups of 2.35× and 2.27× over software-based prefetching and decoupling, respectively. Compared to state-of-the-art hardware, it provides geomean speedups of 1.82× and 1.72× over prefetching and decoupling techniques.
登录
查看更多内容
DOI:
--
发表时间:
1990
期刊:
Twenty-Third Annual Hawaii International Conference on System Sciences
影响因子:
--
作者:
W. Mangione;S. Abraham;E. Davidson
通讯作者:
E. Davidson
影响因子:
1.8
作者:
Kjolstad, Fredrik;Kamil, Shoaib;Amarasinghe, Saman
通讯作者:
Amarasinghe, Saman
DOI:
10.1145/2000064.2000079
发表时间:
2011
期刊:
2011 38th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
作者:
N. Crago;Sanjay J. Patel
通讯作者:
Sanjay J. Patel
DOI:
10.1145/2925426.2926254
发表时间:
2016-06
期刊:
Proceedings of the 2016 International Conference on Supercomputing
影响因子:
--
作者:
S. Ainsworth;Timothy M. Jones
通讯作者:
S. Ainsworth;Timothy M. Jones
DOI:
--
发表时间:
1992
期刊:
[1992] Proceedings the 19th Annual International Symposium on Computer Architecture
影响因子:
--
作者:
W. Wulf
通讯作者:
W. Wulf