Reducing Load Latency with Cache Level Prediction
Reducing Load Latency with Cache Level Prediction
复制标题
DOI:
10.1109/hpca53966.2022.00054
复制
发表时间:
2021-03
期刊:
影响因子:
--
通讯作者:
Majid Jalili;M. Erez
中科院分区:
文献类型:
--
作者:
Majid Jalili;M. Erez
High load latency that results from deep cache hierarchies and relatively slow main memory is an important limiter of single-thread performance. Data prefetch helps reduce this latency by fetching data up the hierarchy before it is requested by load instructions. However, data prefetching has shown to be imperfect in many situations. We propose cache-level prediction to complement prefetchers. Our method predicts which memory hierarchy level a load will access allowing the memory loads to start earlier, and thereby saves many cycles. The predictor provides high prediction accuracy at the cost of just one cycle added latency to L1 misses. Level prediction reduces the memory access latency by 20% on average, and provides speedup of 10.3% over a conventional baseline, and 6.1% over a boosted baseline on generic, graph, and HPC applications.