Feedback Directed Prefetching: Improving the Performance and Bandwidth-Efficiency of Hardware Prefetchers

Feedback Directed Prefetching: Improving the Performance and Bandwidth-Efficiency of Hardware Prefetchers
复制标题

DOI:
10.1109/hpca.2007.346185
复制
发表时间:
2007-02
期刊:
2007 IEEE 13th International Symposium on High Performance Computer Architecture
影响因子:
--
通讯作者:
S. Srinath;O. Mutlu;Hyesoon Kim;Y. Patt
S. Srinath;O. Mutlu;Hyesoon Kim;Y. Patt
中科院分区:
其他
文献类型:
--
作者:
S. Srinath;O. Mutlu;Hyesoon Kim;Y. Patt

文献摘要

被引文献

相似文献

高性能处理器采用硬件数据预取来减少大主存延迟对性能的负面影响。虽然预取在许多程序上大大提高了性能,但它可能会显著降低其他程序的性能。此外,预取会显著增加内存带宽需求。本文提出了一种机制,将动态反馈到预取器的设计,以增加预取提供的性能改善,以及减少预取的负面性能和带宽的影响。我们的机制估计预取器的准确性,预取器的及时性,和预取器造成的缓存污染,以动态调整数据预取器的侵略性。我们介绍了一种新的方法来跟踪缓存污染所造成的预取在运行时。我们还介绍了一种机制,动态地决定在哪里的LRU堆栈中的预取块插入到该高速缓存的基础上造成的该高速缓存污染。与性能最佳的传统基于流的数据预取器配置相比,使用所提出的动态机制在SPEC CPU2000套件中的17个内存密集型基准测试中将平均性能提高了6.5%,同时消耗的内存带宽减少了18.7%。与传统的基于流的数据预取器配置消耗类似的内存带宽量相比,反馈导向的预取提供了13.6%的高性能。实验结果表明,反馈引导预取消除了预取对性能的影响,适用于基于流的预取器、基于全局历史缓冲区的增量相关预取器和基于PC的步长预取器
High performance processors employ hardware data prefetching to reduce the negative performance impact of large main memory latencies. While prefetching improves performance substantially on many programs, it can significantly reduce performance on others. Also, prefetching can significantly increase memory bandwidth requirements. This paper proposes a mechanism that incorporates dynamic feedback into the design of the prefetcher to increase the performance improvement provided by prefetching as well as to reduce the negative performance and bandwidth impact of prefetching. Our mechanism estimates prefetcher accuracy, prefetcher timeliness, and prefetcher-caused cache pollution to adjust the aggressiveness of the data prefetcher dynamically. We introduce a new method to track cache pollution caused by the prefetcher at run-time. We also introduce a mechanism that dynamically decides where in the LRU stack to insert the prefetched blocks in the cache based on the cache pollution caused by the prefetcher. Using the proposed dynamic mechanism improves average performance by 6.5% on 17 memory-intensive benchmarks in the SPEC CPU2000 suite compared to the best-performing conventional stream-based data prefetcher configuration, while it consumes 18.7% less memory bandwidth. Compared to a conventional stream-based data prefetcher configuration that consumes similar amount of memory bandwidth, feedback directed prefetching provides 13.6% higher performance. Our results show that feedback-directed prefetching eliminates the large negative performance impact incurred on some benchmarks due to prefetching, and it is applicable to stream-based prefetchers, global-history-buffer based delta correlation prefetchers, and PC-based stride prefetchers