Prefetching Techniques for Near-memory Throughput Processors

Prefetching Techniques for Near-memory Throughput Processors
复制标题

DOI:
10.1145/2925426.2926282
复制
发表时间:
2016-06
期刊:
Proceedings of the 2016 International Conference on Supercomputing
影响因子:
--
通讯作者:
Reena Panda;Yasuko Eckert;N. Jayasena;Onur Kayiran;Michael Boyer;L. John
Reena Panda;Yasuko Eckert;N. Jayasena;Onur Kayiran;Michael Boyer;L. John
中科院分区:
其他
文献类型:
--
作者:
Reena Panda;Yasuko Eckert;N. Jayasena;Onur Kayiran;Michael Boyer;L. John

文献摘要

被引文献

相似文献

最近,近序列处理或内存处理(PIM)最近重新引起了很多兴趣,作为一个可行的解决方案,以克服记忆墙所面临的挑战。这种趋势主要是由于3D堆积的记忆的出现而推动了这种趋势。由于其出色的带宽利用能力,GPU被吹捧为内存处理器的出色候选者。尽管将GPU核心放在记忆下,但在本文中将其暴露于前所未有的内存带宽,但我们证明,仍然存在重要的机会,可以通过改善其内存性能来改善更简单的内存GPU处理器(GPU-PIM)的性能。因此,我们提出了三个轻量级,实用的记忆侧预脱毛,以提高GPU-PIM系统的性能。拟议的预摘要利用单个内存访问中的模式和波前 - 定位的内存流中的协同作用,结合对内存系统状态的更好理解,从DRAM行缓冲区预取到片上预取缓冲区,从而实现了75%以上预取准精度和行缓冲区的40%提高。为了最大程度地利用预取的数据并最大程度地减少thrashing,预摘要还基于独特的死行预测机制以及基于驱逐的预取触发策略来控制其侵略性。与基线相比,拟议的预赠果会提高超过60%(最大)和平均9%的绩效,同时使用少于5.6kb的其他硬件获得了Perfect-L2的33%以上的性能优势。拟议的预科师还优于最先进的内存侧预拿,OWL超过20%。
Near-memory processing or processing-in-memory (PIM) is regaining a lot of interest recently as a viable solution to overcome the challenges imposed by memory wall. This trend has been mainly fueled by the emergence of 3D-stacked memories. GPUs are touted as great candidates for in-memory processors due to their superior bandwidth utilization capabilities. Although putting a GPU core beneath memory exposes it to unprecedented memory bandwidth, in this paper, we demonstrate that significant opportunities still exist to improve the performance of the simpler, in-memory GPU processors (GPU-PIM) by improving their memory performance. Thus, we propose three light-weight, practical memory-side prefetchers to improve the performance of GPU-PIM systems. The proposed prefetchers exploit the patterns in individual memory accesses and synergy in the wavefront-localized memory streams, combined with a better understanding of the memory-system state, to prefetch from DRAM row buffers into on-chip prefetch buffers, thereby achieving over 75% prefetcher accuracy and 40% improvement in row buffer locality. In order to maximize utilization of prefetched data and minimize thrashing, the prefetchers also use a novel prefetch buffer management policy based on a unique dead-row prediction mechanism together with an eviction-based prefetch-trigger policy to control their aggressiveness. The proposed prefetchers improve performance by over 60% (max) and 9% on average as compared to the baseline, while achieving over 33% of the performance benefits of perfect-L2 using less than 5.6KB of additional hardware. The proposed prefetchers also outperform the state-of-the-art memory-side prefetcher, OWL by more than 20%.