A Case for Resource Efficient Prefetching in Multicores

A Case for Resource Efficient Prefetching in Multicores
复制标题

多核资源高效预取的案例

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Parallel Processing
影响因子:
--
通讯作者:
Erik Hagersten
Erik Hagersten
中科院分区:
--
文献类型:
--
作者:
Muneeb Khan;Andreas Sandberg;Erik Hagersten

文献摘要

被引文献

相似文献

现代处理器通常采用复杂的预取技术来隐藏内存延迟。硬件预取的事实证明非常有效,并且在隔离运行时可以加快某些规格CPU 2006基准的40%以上。但是,这种速度通常是以预购大量无用的数据(有时是所需数据的两倍以上)为代价的,这些数据浪费了最后一个级别的缓存空间和芯片外带宽。本文探讨了准确的资源效率预取方案如何通过保存多学院中的共享资源来使性能受益。我们提出了一个使用低空的运行时采样和快速缓存建模的框架,以准确识别经常错过缓存中的内存指令。然后,我们使用此信息自动将软件预插入在应用程序中。我们的预取方案具有良好的准确性,并尽可能采用缓存。这些属性有助于减少芯片外带宽消耗和最后一级的缓存污染。虽然单线程的性能仍然与硬件预取相匹配,但是当使用多个核心并要求共享资源的需求增长时,该方案的全部优势将实现。我们评估了两个现代商品多头的方法。在180种完全利用多功能的混合工作负载中,所提出的软件预取机制的吞吐量比硬件预取的要好得多24%,并且平均表现要高10%。
Modern processors typically employ sophisticated prefetching techniques for hiding memory latency. Hardware prefetching has proven very effective and can speed up some SPEC CPU 2006 benchmarks by more than 40% when running in isolation. However, this speedup often comes at the cost of prefetching a significant volume of useless data (sometimes more than twice the data required) which wastes shared last level cache space and off-chip bandwidth. This paper explores how an accurate resource-efficient prefetching scheme can benefit performance by conserving shared resources in multicores. We present a framework that uses low-overhead runtime sampling and fast cache modeling to accurately identify memory instructions that frequently miss in the cache. We then use this information to automatically insert software prefetches in the application. Our prefetching scheme has good accuracy and employs cache bypassing whenever possible. These properties help reduce off-chip bandwidth consumption and last-level cache pollution. While single-thread performance remains comparable to hardware prefetching, the full advantage of the scheme is realized when several cores are used and demand for shared resources grows. We evaluate our method on two modern commodity multicores. Across 180 mixed workloads that fully utilize a multicore, the proposed software prefetching mechanism achieves up to 24% better throughput than hardware prefetching, and performs 10% better on average.