Vector Runahead

Vector Runahead
复制标题

矢量奔跑

DOI:
--
复制
发表时间:
2021
期刊:
International Symposium on Computer Architecture
影响因子:
--
通讯作者:
L. Eeckhout
L. Eeckhout
中科院分区:
--
文献类型:
--
作者:
Ajeya Naithani;S. Ainsworth;Timothy M. Jones;L. Eeckhout

文献摘要

参考文献

被引文献

相似文献

内存墙对许多现代工作负载的性能造成了很大的限制。这些应用程序具有复杂的依赖、间接内存访问链,即使是最先进的微架构预取器也无法拾取。其结果是,当前的无序超标量处理器大部分时间都处于停滞状态。虽然有可能构建特殊用途的体系结构来利用基本的内存级并行性,但在传统处理器中自动提高其性能的微体系结构技术仍然难以实现。提前执行对于隐藏程序执行中的延迟是一个诱人的提议。但是,为了实现高内存级并行性,标准的提前执行会跳过缓存丢失。在现代工作负载中,这意味着它只预取每个依赖链中的第一个缓存丢失负载。我们认为这不是一个基本的限制。如果提前运行在缓存丢失时停止以生成依赖的链负载,那么如果它可以一次在多个链上停止,则可以恢复性能。有了这一见解,我们提出了Vector Runahead,这是一种技术,它可以预取整个负载链,并推测地将多个循环迭代中的标量操作重新排序为矢量格式,从而一次引入许多独立的负载。向前运行指令流的矢量化增加了有效的读取/解码带宽,同时减少了资源需求,以更快的速度实现了高度的内存级并行性。在各种内存延迟受限的间接工作负载中,Vector Runahead在大型无序超标量系统上实现了1.79倍的性能加速,显著改进了最先进的技术。
The memory wall places a significant limit on performance for many modern workloads. These applications feature complex chains of dependent, indirect memory accesses, which cannot be picked up by even the most advanced microar-chitectural prefetchers. The result is that current out-of-order superscalar processors spend the majority of their time stalled. While it is possible to build special-purpose architectures to exploit the fundamental memory-level parallelism, a microarchi-tectural technique to automatically improve their performance in conventional processors has remained elusive.Runahead execution is a tempting proposition for hiding latency in program execution. However, to achieve high memory-level parallelism, a standard runahead execution skips ahead of cache misses. In modern workloads, this means it only prefetches the first cache-missing load in each dependent chain. We argue that this is not a fundamental limitation. If runahead were instead to stall on cache misses to generate dependent chain loads, then it could regain performance if it could stall on many at once. With this insight, we present Vector Runahead, a technique that prefetches entire load chains and speculatively reorders scalar operations from multiple loop iterations into vector format to bring in many independent loads at once. Vectorization of the runahead instruction stream increases the effective fetch/decode bandwidth with reduced resource requirements, to achieve high degrees of memory-level parallelism at a much faster rate. Across a variety of memory-latency-bound indirect workloads, Vector Runahead achieves a 1.79× performance speedup on a large out-of-order superscalar system, significantly improving on state-of-the-art techniques.
DOI: 10.1145/3373376.3378498
发表时间: 2020-03
期刊: Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子: --
作者:
Grant Ayers;Heiner Litz;Christos Kozyrakis;Parthasarathy Ranganathan
通讯作者: Grant Ayers;Heiner Litz;Christos Kozyrakis;Parthasarathy Ranganathan
DOI: 10.1145/2925426.2926254
发表时间: 2016-06
期刊: Proceedings of the 2016 International Conference on Supercomputing
影响因子: --
作者:
S. Ainsworth;Timothy M. Jones
通讯作者: S. Ainsworth;Timothy M. Jones
针对不规则工作负载的事件触发可编程预取器
DOI: 10.1145/3173162.3173189
发表时间: 2018
期刊: --
影响因子: --
作者:
Ainsworth S
通讯作者: Ainsworth S
限制自动矢量化:少即是多
DOI: 10.1109/pact.2015.32
发表时间: 2015
期刊: --
影响因子: --
作者:
Porpodas V
通讯作者: Porpodas V