Reducing Load Latency with Cache Level Prediction

Reducing Load Latency with Cache Level Prediction
复制标题

DOI:
10.1109/hpca53966.2022.00054
复制
发表时间:
2021-03
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Majid Jalili;M. Erez
Majid Jalili;M. Erez
中科院分区:
其他
文献类型:
--
作者:
Majid Jalili;M. Erez

文献摘要

相似文献

深缓存层次结构和相对较慢的主内存导致的高负载延迟是单线程性能的重要限制因素。数据预取通过在加载指令请求数据之前沿层次结构向上获取数据来帮助减少此延迟。然而,数据预取在许多情况下是不完美的。我们提出缓存级预测来补充预取器。我们的方法预测的内存层次结构级别的负载将访问允许内存加载开始较早,从而节省了许多周期。预测器提供了高预测精度,代价是L1未命中仅增加了一个周期的延迟。级别预测将内存访问延迟平均降低了20%,并在常规基线上提供了10.3%的加速,在通用、图形和HPC应用程序上提供了6.1%的加速。
High load latency that results from deep cache hierarchies and relatively slow main memory is an important limiter of single-thread performance. Data prefetch helps reduce this latency by fetching data up the hierarchy before it is requested by load instructions. However, data prefetching has shown to be imperfect in many situations. We propose cache-level prediction to complement prefetchers. Our method predicts which memory hierarchy level a load will access allowing the memory loads to start earlier, and thereby saves many cycles. The predictor provides high prediction accuracy at the cost of just one cycle added latency to L1 misses. Level prediction reduces the memory access latency by 20% on average, and provides speedup of 10.3% over a conventional baseline, and 6.1% over a boosted baseline on generic, graph, and HPC applications.