Exploring Predictive Replacement Policies for Instruction Cache and Branch Target Buffer

Exploring Predictive Replacement Policies for Instruction Cache and Branch Target Buffer
复制标题

探索指令缓存和分支目标缓冲区的预测替换策略

DOI:
10.1109/isca.2018.00050
复制
发表时间:
2018
期刊:
Proceedings of the 45th Annual International Symposium on Computer Architecture
影响因子:
--
通讯作者:
Jimenez, Daniel A.
Jimenez, Daniel A.
中科院分区:
--
文献类型:
--
作者:
Mirbagher Ajorpaz, Samira;Garza, Elba;Jindal, Sangam;Jimenez, Daniel A.

文献摘要

参考文献

被引文献

相似文献

现代处理器支持使用指令高速缓存 (I-cache) 和分支目标缓冲区 (BTB) 进行指令提取。由于时序和面积限制,I-cache 和 BTB 必须有效地利用其有限的容量。 I-cache 中的块或 BTB 中重用潜力较低的条目应替换为更有用的块/条目。这项工作探索了基于重用预测的预测替换策略,该策略可应用于 I-cache 和 BTB。使用大量最近发布的工业跟踪数据,我们表明预测性替换策略可以减少 I-cache 和 BTB 中的缺失。我们引入全局历史重用预测 (GHRP),这是一种替换技术,它使用过去指令地址的历史及其重用行为来预测 I-cache 中的死块和 BTB 中的死条目。本文描述了 GHRP 作为 I-cache 和 BTB 的死块替换和旁路优化的有效性。对于具有 64B 块大小的 64KB 组关联 I-cache,GHRP 与一组 662 个工业工作负载上的最近最少使用 (LRU) 策略相比,将每 1000 条指令 (MPKI) 的 I-cache 未命中率平均降低了 18%,其性能明显优于静态重新引用间隔预测 (SRRIP) 和采样死块预测 (SDBP)。对于 4K 条目 BTB,GHRP 比 LRU 平均降低 MPKI 30%,比 SRRIP 平均降低 23%,比 SDBP 平均降低 29%。
Modern processors support instruction fetch with the instruction cache (I-cache) and branch target buffer (BTB). Due to timing and area constraints, the I-cache and BTB must efficiently make use of their limited capacities. Blocks in the I-cache or entries in the BTB that have low potential for reuse should be replaced by more useful blocks/entries. This work explores predictive replacement policies based on reuse prediction that can be applied to both the I-cache and BTB. Using a large suite of recently released industrial traces, we show that predictive replacement policies can reduce misses in the I-cache and BTB. We introduce Global History Reuse Prediction (GHRP), a replacement technique that uses the history of past instruction addresses and their reuse behaviors to predict dead blocks in the I-cache and dead entries in the BTB. This paper describes the effectiveness of GHRP as a dead block replacement and bypass optimization for both the I-cache and BTB. For a 64KB set-associative I-cache with a 64B block size, GHRP lowers the I-cache misses per 1000 instructions (MPKI) by an average of 18% over the least-recently-used (LRU) policy on a set of 662 industrial workloads, performing significantly better than Static Re-reference Interval Prediction (SRRIP) and Sampling Dead Block Prediction (SDBP). For a 4K-entry BTB, GHRP lowers MPKI by an average of 30% over LRU, 23% over SRRIP, and 29% over SDBP.
分支目标缓冲区设计与优化
DOI: 10.1109/12.214687
发表时间: 1993
期刊: IEEE Trans. Computers
影响因子: --
作者:
Chris H. Perleberg;A. Smith
通讯作者: A. Smith
DOI: --
发表时间: 1980
影响因子: 3.7
作者:
R. Holgate;R. Ibbett
通讯作者: R. Ibbett
两级批量预加载分支预测
DOI: 10.1109/hpca.2013.6522308
发表时间: 2013
期刊: 2013 IEEE 19th International Symposium on High Performance Computer Architecture (HPCA)
影响因子: --
作者:
J. Bonanno;Adam Collura;Daniel Lipetz;U. Mayer;Brian R. Prasky;A. Saporito
通讯作者: A. Saporito
Boomerang:用于控制流交付的无元数据架构
DOI: 10.1109/hpca.2017.53
发表时间: 2017
期刊: 2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子: --
作者:
Rakesh Kumar;Cheng;Boris Grot;V. Nagarajan
通讯作者: V. Nagarajan
延迟对分支预测器设计的影响
DOI: --
发表时间: 2000
期刊: Proceedings 33rd Annual IEEE/ACM International Symposium on Microarchitecture. MICRO-33 2000
影响因子: --
作者:
Daniel A. Jiménez;S. Keckler;Calvin Lin
通讯作者: Calvin Lin