Energy-Efficient GPU L2 Cache Design Using Instruction-Level Data Locality Similarity

Energy-Efficient GPU L2 Cache Design Using Instruction-Level Data Locality Similarity
复制标题

使用指令级数据局部性相似性的节能 GPU L2 缓存设计

DOI:
10.1145/3408060
复制
发表时间:
2020
影响因子:
1.4
通讯作者:
Xin Fu
Xin Fu
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jingweijia Tan;Kaige Yan;Shuaiwen Leon Song;Xin Fu

文献摘要

相似文献

本文提出了一种适用于大规模并行、面向吞吐量的架构(如GPU)的新型高能效缓存设计。与现代GPU上的L1数据缓存不同,所有流多处理器共享的L2缓存并不是主要的性能瓶颈,但它确实消耗了大量的芯片能量。我们观察到,L2缓存的利用率很低,95.6%的时间用于存储无用数据。如果识别并减少L2上的这种“死时间”,L2‘S的能效将会大幅提高。幸运的是,我们发现GPU的SIMT编程模型在线程之间提供了一个独特的特征:指令级数据局部性相似性,该相似性可以用于准确预测L2缓存块级别的数据重新引用计数。我们提出了一个简单的设计,利用这种局部相似性来构建一个节能的GPU L2Cache,命名为LoSCache。具体来说,LoSCache使用来自一小组协作线程数组的数据局部性信息来动态预测剩余协作线程数组的L2级数据引用计数。在此之后,如果特定的L2高速缓存线在某些访问之后被预测为“死”,则它们可以被断电。实验结果表明,在性能损失仅为0.5%的情况下,我们提出的设计方案可以显著降低L2缓存的能量,平均降低%。此外,LoSCache具有成本效益,独立于调度策略,并与最先进的一级缓存设计兼容,以实现额外的能源节约。
This article presents a novel energy-efficient cache design for massively parallel, throughput-oriented architectures like GPUs. Unlike L1 data cache on modern GPUs, L2 cache shared by all of the streaming multiprocessors is not the primary performance bottleneck, but it does consume a large amount of chip energy. We observe that L2 cache is significantly underutilized by spending 95.6% of the time storing useless data. If such “dead time” on L2 is identified and reduced, L2’s energy efficiency can be drastically improved. Fortunately, we discover that the SIMT programming model of GPUs provides a unique feature among threads: instruction-level data locality similarity, which can be used to accurately predict the data re-reference counts at L2 cache block level. We propose a simple design that leverages thisLocalitySimilarity to build an energy-efficient GPU L2Cache, namedLoSCache. Specifically, LoSCache uses the data locality information from a small group of cooperative thread arrays to dynamically predict the L2-level data re-reference counts of the remaining cooperative thread arrays. After that, specific L2 cache lines can be powered off if they are predicted to be “dead” after certain accesses. Experimental results on a wide range of applications demonstrate that our proposed design can significantly reduce the L2 cache energy by an average of 64% with only 0.5% performance loss. In addition, LoSCache is cost effective, independent of the scheduling policies, and compatible with the state-of-the-art L1 cache designs for additional energy savings.