Stream Floating: Enabling Proactive and Decentralized Cache Optimizations

Stream Floating: Enabling Proactive and Decentralized Cache Optimizations
复制标题

DOI:
10.1109/hpca51647.2021.00060
复制
发表时间:
2021-02
期刊:
2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Zhengrong Wang;Jian Weng;Jason Lowe-Power;Jayesh Gaur;Tony Nowatzki
Zhengrong Wang;Jian Weng;Jason Lowe-Power;Jayesh Gaur;Tony Nowatzki
中科院分区:
其他
文献类型:
--
作者:
Zhengrong Wang;Jian Weng;Jason Lowe-Power;Jayesh Gaur;Tony Nowatzki

文献摘要

被引文献

相似文献

随着多核心系统的规模和芯片内存储能力继续增长,片上网络带宽和潜伏期成为有问题的瓶颈。因此,数据传输的开销,相干协议和替换策略变得越来越重要。不幸的是,即使在结构良好的程序中,由于传统的高速缓存层次结构的反应性和集中性质,许多自然优化也很难实施,在短,核心启动所有请求中,所有请求均可用于简短的高速缓存线粒状访问。例如,可以从共享的缓存中流出持久的访问模式,而无需来自核心的请求。可以通过从缓存内部制成的链接请求来执行间接内存访问,而不是不断返回核心。我们的主要见解是,如果程序可以嵌入有关其ISA中的长期内存流行为的信息,那么这些流可以将其浮动到适当的内存层次结构级别。解决生成和缓存请求的分散方法可以通过在内核请求之前主动发送数据来导致更好的缓存策略,并降低请求和数据流量。为了评估流浮动的机会,我们通过流动机增强了一个瓷砖的多层缓存层次结构,以处理最后一级的高速缓存库中的流请求。我们开发了几种新颖的优化,这些优化是通过ISA中的流暴露促进的,随后暴露于缓存。我们使用基于周期的执行驱动的GEM5模拟器进行评估,使用Rodinia的10个数据处理工作负载和2个用OpenMP编写的流媒体内核。我们发现,溪流浮动能够分别在井井有条和OOO核心上进行52%和39%的速度,分别具有最先进的预摘要设计,具有64%和49%的能源效率优势。
As multicore systems continue to grow in scale and on-chip memory capacity, the on-chip network bandwidth and latency become problematic bottlenecks. Because of this, overheads in data transfer, the coherence protocol and replacement policies become increasingly important. Unfortunately, even in well-structured programs, many natural optimizations are difficult to implement because of the reactive and centralized nature of traditional cache hierarchies, where all requests are initiated by the core for short, cache line granularity accesses. For example, long-lasting access patterns could be streamed from shared caches without requests from the core. Indirect memory access can be performed by chaining requests made from within the cache, rather than constantly returning to the core. Our primary insight is that if programs can embed information about long-term memory stream behavior in their ISAs, then these streams can be floated to the appropriate level of the memory hierarchy. This decentralized approach to address generation and cache requests can lead to better cache policies and lower request and data traffic by proactively sending data before the cores even request it. To evaluate the opportunities of stream floating, we enhance a tiled multicore cache hierarchy with stream engines to process stream requests in last-level cache banks. We develop several novel optimizations that are facilitated by stream exposure in the ISA, and subsequent exposure to caches. We evaluate using a cycle-level execution-driven gem5-based simulator, using 10 data-processing workloads from Rodinia and 2 streaming kernels written in OpenMP. We find that stream floating enables 52% and 39% speedup over an inorder and OOO core with state of art prefetcher design respectively, with 64% and 49% energy efficiency advantage.