Stealth prefetching

Stealth prefetching
复制标题

隐形预取

DOI:
10.1145/1168857.1168892
复制
发表时间:
2006
期刊:
Proceedings of the 26th International Symposium on Computer Architecture (Cat. No.99CB36367)
影响因子:
--
通讯作者:
James E. Smith
James E. Smith
中科院分区:
--
文献类型:
--
作者:
J. F. Cantin;Mikko H. Lipasti;James E. Smith

文献摘要

被引文献

相似文献

在共享内存的多处理器系统中,预取是一个日益困难的问题。随着系统设计逐渐包含更多更快的处理器,内存延迟和互连流量会增加。虽然主动预取技术可以缓解不断增加的存储延迟,但它们会浪费宝贵的互连带宽和过早访问共享数据,导致远程节点的状态降级,从而迫使以后的升级。隐形预取利用CGCT来识别未被其他处理器共享的内存区域,以打开页面模式从DRAM中积极地获取这些行,并在预期未来引用的情况下将它们移至接近处理器的位置。我们对商业、科学和多程序工作负载的分析表明,与使用常规预取的激进基准系统相比,隐形预取平均加速20%。
Prefetching in shared-memory multiprocessor systems is an increasingly difficult problem. As system designs grow to incorporate larger numbers of faster processors, memory latency and interconnect traffic increase. While aggressive prefetching techniques can mitigate the increasing memory latency, they can harm performance by wasting precious interconnect bandwidth and prematurely accessing shared data, causing state downgrades at remote nodes that force later upgrades.This paper investigates Stealth Prefetching, a new technique that utilizes information from Coarse-Grain Coherence Tracking (CGCT) for prefetching data aggressively, stealthily, and efficiently in a broadcast-based shared-memory multiprocessor system. Stealth Prefetching utilizes CGCT to identify regions of memory that are not shared by other processors, aggressively fetches these lines from DRAM in open-page mode, and moves them close to the processor in anticipation of future references. Our analysis with commercial, scientific, and multiprogrammed workloads show that Stealth Prefetching provides an average speedup of 20% over an aggressive baseline system with conventional prefetching.