Stream-based Memory Access Specialization for General Purpose Processors

Stream-based Memory Access Specialization for General Purpose Processors
复制标题

DOI:
10.1145/3307650.3322229
复制
发表时间:
2019-06
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
Zhengrong Wang;Tony Nowatzki
Zhengrong Wang;Tony Nowatzki
中科院分区:
其他
文献类型:
--
作者:
Zhengrong Wang;Tony Nowatzki

文献摘要

相似文献

由于技术扩展的严重限制,架构师已经在专门用于计算原语(例如向量指令,循环加速器)的通用处理器方面进行了创新。一般原则是向伊萨公开丰富的语义。一个探索的机会是是否更丰富的语义的内存访问模式也可以用来提高内存和通信的效率。两个重要的开放问题是如何传达更高级别的内存信息以及如何在硬件中利用这些信息。我们发现,大多数的内存访问遵循少量的简单模式,我们这些流(如仿射,间接)。流通常可以从核心执行中解耦,并且模式可以持续足够长的时间来表达有用的行为。因此,我们的方法是将流表示为伊萨原语,我们认为可以实现:预取流访问以隐藏内存延迟,半绑定解耦访问以消除地址计算并优化内存接口,并最终通知缓存策略。在这项工作中,我们提出了ISA扩展的并行流,与核心使用FIFO为基础的接口。我们在积极的广泛发行OOO内核上为上述每个机会实施优化,并使用SPEC CPU 2017和CortexSuite进行评估[1],[2]。在所有工作负载中,我们观察到硬件步幅预取约1.37倍的加速和能效提升。
Because of severe limitations in technology scaling, architects have innovated in specializing general purpose processors for computation primitives (e.g. vector instructions, loop accelerators). The general principle is exposing rich semantics to the ISA. An opportunity to explore is whether richer semantics of memory access patterns could also be used to improve the efficiency of memory and communication. Two important open questions are how to convey higher level memory information and how to take advantage of this information in hardware. We find that a majority of memory accesses follow a small number of simple patterns; we term these streams (e.g. affine, indirect). Streams can often be decoupled from core execution, and patterns persist long enough to express useful behavior. Our approach is therefore to express streams as ISA primitives, which we argue can enable: prefetch stream accesses to hide memory latency, semi- binding decoupled access to remove address computation and optimize the memory interface, and finally inform cache policies. In this work, we propose ISA-extensions for decoupled-streams, which interact with the core using a FIFO-based interface. We implement optimizations for each of the aforementioned opportunities on an aggressive wide-issue OOO core and evaluate with SPEC CPU 2017 and CortexSuite[1], [2]. Across all workloads, we observe about 1.37x speedup and energy efficiency improvement over hardware stride prefetching.