ContextPreRF: Enhancing the Performance and Energy of GPUs With Nonuniform Register Access

ContextPreRF: Enhancing the Performance and Energy of GPUs With Nonuniform Register Access
复制标题

ContextPreRF:通过非统一寄存器访问增强 GPU 的性能和能耗

DOI:
10.1109/tvlsi.2015.2397876
复制
发表时间:
2016
影响因子:
2.8
通讯作者:
A. Jones
A. Jones
中科院分区:
工程技术2区
文献类型:
--
作者:
Michael Moeng;Haifeng Xu;R. Melhem;A. Jones

文献摘要

被引文献

相似文献

寄存器堆是影响图形处理单元(GPU)指令吞吐量的关键数据存储单元。通常,GPU寄存器堆非常大,以容纳许多并发线程,并且使用与片上高速缓存相同的SRAM技术来实现。本文提出了一种新的寄存器堆结构ContExtrf,它有效地利用了具有非均匀访问特性的寄存器堆,包括混合静态随机存取存储器/动态随机存取存储器(S/D)和自旋电子磁区壁存储器(DEM)。ContExtrf允许在GPU内的同一区域实现更大容量的寄存器堆,同时降低功耗。我们还提出了一种隐藏切换延迟的硬件预切换方案--一旦寄存器请求排队,包含相应寄存器的非均匀访问存储器就被发送抢占式切换请求。因此,我们的方案透明地隐藏了在寄存器上下文之间切换的惩罚。用S/D代替寄存器堆静态存储器后,功耗降低了37%,平均性能下降了1.4%。使用DWM,我们将寄存器堆的能量减少了74%,平均性能损失为0.4%。对于密度更高的DWM,我们模拟将节省的区域转换为额外的寄存器、高速缓存和共享内存-这比基准SRAM寄存器堆的性能提高了13.5%。
Register files are a key data storage unit that impacts instruction throughput for graphics processing units (GPUs). Typically, GPU register files are quite large to accommodate many concurrent threads and are implemented using the same SRAM technology as the on-chip cache. We propose contextrf, a new register file architecture that efficiently leverages register files with nonuniform access characteristics, including hybrid SRAM/DRAM (S/D) and spintronic domain-wall memories (DWMs). Contextrf allows greater-capacity register files to be implemented in the same area within the GPU, with reduced power consumption. We also propose contextPreRF, a hardware preswitch scheme to hide switching delays-as soon as a register request is queued, the nonuniform access memories containing the corresponding register are sent a preemptive switch request. Thus, our scheme transparently hides the penalties of switching between register contexts. After replacing the register file SRAM with S/D, we can reduce energy by 37%, with a 1.4% average performance drop. Employing DWM, we reduce register file energy by 74%, with a 0.4% average performance penalty. For the denser DWM, we model converting the saved area into additional registers, cache, and shared memory-this improves performance by 13.5% over the baseline SRAM register file.