Kill the Program Counter: Reconstructing Program Behavior in the Processor Cache Hierarchy

Kill the Program Counter: Reconstructing Program Behavior in the Processor Cache Hierarchy
复制标题

杀死程序计数器:重建处理器缓存层次结构中的程序行为

DOI:
10.1145/3037697.3037701
复制
发表时间:
2017
期刊:
Proceedings of the Twenty-Second International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
C. Wilkerson
C. Wilkerson
中科院分区:
--
文献类型:
--
作者:
Jinchun Kim;Elvira Teran;Paul V. Gratz;Daniel A. Jiménez;Seth H. Pugsley;C. Wilkerson

文献摘要

参考文献

被引文献

相似文献

在高性能微处理器的设计中,数据预取和缓存替换算法一直是研究的热点。通常,数据预取器在私有高速缓存中操作,并且不与共享末级高速缓存(LLC)中的替换策略交互。类似地,大多数替换策略不将需求和预取请求视为不同类型的请求。特别地,基于程序计数器(PC)的替换策略不能从预取请求中学习,因为数据预取器不生成PC值。基于PC的策略也可能受到编译器优化的负面影响。在本文中,我们提出了一个整体的缓存管理技术,称为杀死PC(KPC),克服了传统的预取和替换策略算法的弱点。KPC缓存管理有三个新的贡献。首先,预取器基于其预测置信度来近似预取请求的未来使用距离。第二,简单的替换策略提供了与使用全局滞后的当前最先进的基于PC的预测相似或更好的性能。第三,KPC将预取和替换策略集成到一个整体系统中,其整体性能大于各部分性能之和。来自预取器的信息用于提高替换策略的性能,反之亦然。最后,KPC消除了通过整个片上高速缓存层次结构传播PC的需要,同时提供了一个整体的高速缓存管理方法,其性能优于最先进的基于PC和非基于PC的方案。我们的评估表明,KPC提供了8%以上的性能比现有的预取器和多核工作负载的替换策略的最佳组合。
Data prefetching and cache replacement algorithms have been intensively studied in the design of high performance microprocessors. Typically, the data prefetcher operates in the private caches and does not interact with the replacement policy in the shared Last-Level Cache (LLC). Similarly, most replacement policies do not consider demand and prefetch requests as different types of requests. In particular, program counter (PC)-based replacement policies cannot learn from prefetch requests since the data prefetcher does not generate a PC value. PC-based policies can also be negatively affected by compiler optimizations. In this paper, we propose a holistic cache management technique called Kill-the-PC (KPC) that overcomes the weaknesses of traditional prefetching and replacement policy algorithms. KPC cache management has three novel contributions. First, a prefetcher which approximates the future use distance of prefetch requests based on its prediction confidence. Second, a simple replacement policy provides similar or better performance than current state-of-the-art PC-based prediction using global hysteresis. Third, KPC integrates prefetching and replacement policy into a whole system which is greater than the sum of its parts. Information from the prefetcher is used to improve the performance of the replacement policy and vice-versa. Finally, KPC removes the need to propagate the PC through entire on-chip cache hierarchy while providing a holistic cache management approach with better performance than state-of-the-art PC-, and non-PC-based schemes. Our evaluation shows that KPC provides 8% better performance than the best combination of existing prefetcher and replacement policy for multi-core workloads.
用于重用预测的感知器学习
DOI: 10.1109/micro.2016.7783705
发表时间: 2016
期刊: 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO
影响因子: --
作者:
Teran, Elvira;Wang, Zhe;Jimenez, Daniel A.
通讯作者: Jimenez, Daniel A.