Designing a Modern Memory Hierarchy with Hardware Prefetching

Designing a Modern Memory Hierarchy with Hardware Prefetching
复制标题

DOI:
10.1109/12.966495
复制
发表时间:
2001-11
期刊:
IEEE Trans. Computers
影响因子:
--
通讯作者:
Wei-Fen Lin;S. Reinhardt;D. Burger
Wei-Fen Lin;S. Reinhardt;D. Burger
中科院分区:
其他
文献类型:
--
作者:
Wei-Fen Lin;S. Reinhardt;D. Burger

文献摘要

被引文献

相似文献

在本文中,我们解决了严重的性能差距所造成的高处理器时钟速率和缓慢的DRAM访问。我们表明,即使有一个积极的,下一代的内存系统,使用四个直接Rambus通道和集成的一兆字节的二级缓存,处理器仍然花了一半以上的时间停在L2的失误。我们的实验分析开始于积极调整我们的基线存储器系统的努力:结合优化,以减少DRAM行缓冲器未命中,重新排序未命中访问,以减少排队延迟,并调整L2块大小,以匹配每个通道组织。我们发现,有一个很大的差距,在性能最好的块大小,并在其中错过率最小化。使用这些结果,我们评估的硬件预取单元集成的L2高速缓存和内存控制器。通过仅在Rambus通道空闲时发出预取,优先考虑它们以最大化DRAM行缓冲器命中,并给予它们低替换优先级,我们在26个SPEC2000基准测试中的10个上实现了65%的加速,而不会降低其他测试的性能。有了8个Rambus通道,这10个基准测试的性能提高到完美的L2缓存的10%以内。
In this paper, we address the severe performance gap caused by high processor clock rates and slow DRAM accesses. We show that, even with an aggressive, next-generation memory system using four Direct Rambus channels and an integrated one-megabyte level-two cache, a processor still spends over half its time stalling for L2 misses. Our experimental analysis begins with an effort to tune our baseline memory system aggressively: incorporating optimizations to reduce DRAM row buffer misses, reordering miss accesses to reduce queuing delay, and adjusting the L2 block size to match each channel organization. We show that there is a large gap between the block sizes at which performance is best and at which miss rate is minimized. Using those results, we evaluate a hardware prefetch unit integrated with the L2 cache and memory controllers. By issuing prefetches only when the Rambus channels are idle, prioritizing them to maximize DRAM row buffer hits, and giving them low replacement priority, we achieve a 65 percent speedup across 10 of the 26 SPEC2000 benchmarks, without degrading the performance of the others. With eight Rambus channels, these 10 benchmarks improve to within 10 percent of the performance of a perfect L2 cache.