Shielding STT-RAM Based Register Files on GPUs against Read Disturbance

Shielding STT-RAM Based Register Files on GPUs against Read Disturbance
复制标题

DOI:
10.1145/2996191
复制
发表时间:
2016-11
期刊:
ACM Journal on Emerging Technologies in Computing Systems (JETC)
影响因子:
--
通讯作者:
Hang Zhang;Xuhao Chen;Nong Xiao;Lei Wang;Fang Liu;Wei Chen;Zhiguang Chen
Hang Zhang;Xuhao Chen;Nong Xiao;Lei Wang;Fang Liu;Wei Chen;Zhiguang Chen
中科院分区:
其他
文献类型:
--
作者:
Hang Zhang;Xuhao Chen;Nong Xiao;Lei Wang;Fang Liu;Wei Chen;Zhiguang Chen

文献摘要

被引文献

相似文献

为了解决GPU上SRAM的高能耗问题,新兴的自旋转移力矩(STT-RAM)存储器技术已经被深入研究以构建GPU寄存器文件以获得更好的能效,这要归功于其低泄漏功率、高密度和良好的可扩展性的优点。然而,STT-RAM遭受读干扰问题,这是由于读电流和写电流之间的电压差随着技术规模而变小的事实。读干扰导致读操作的高错误率,这不能被GPU的大容量寄存器文件上的SEC-DED ECC有效地保护。现有方案(例如,读取-恢复)来减轻读取干扰通常会导致重要的性能损失或过多的能量开销,因此不适用于旨在实现高性能和能量效率的GPU寄存器文件设计。为了对抗读干扰,我们提出了一种新颖的软硬件协同设计的解决方案(即,Red-Shield),它包括三个优化,以克服现有解决方案的局限性。首先,我们在编译阶段识别死读并增加指令以避免不必要的恢复。其次,我们采用一个小的读缓冲区,以适应寄存器读取高访问局部性,以进一步减少恢复。第三,我们提出了一种自适应恢复机制,根据相应寄存器组的忙碌状态选择合适的恢复方案。实验结果表明,我们提出的设计可以有效地减轻性能损失和能量开销所造成的恢复操作,同时仍然保持读取的可靠性。
To address the high energy consumption issue of SRAM on GPUs, emerging Spin-Transfer Torque (STT-RAM) memory technology has been intensively studied to build GPU register files for better energy-efficiency, thanks to its benefits of low leakage power, high density, and good scalability. However, STT-RAM suffers from the read disturbance issue, which stems from the fact that the voltage difference between read current and write current becomes smaller as technology scales. The read disturbance leads to high error rates for read operations, which cannot be effectively protected by the SEC-DED ECC on large-capacity register files of GPUs. Prior schemes (e.g., read-restore) to mitigate the read disturbance usually incur either non-trivial performance loss or excessive energy overhead, thus not applicable for the GPU register file design that aims to achieve both high performance and energy-efficiency. To combat the read disturbance, we propose a novel software-hardware co-designed solution (i.e., Red-Shield), which consists of three optimizations to overcome the limitations of the existing solutions. First, we identify dead reads at compiling stage and augment instructions to avoid unnecessary restores. Second, we employ a small read buffer to accommodate register reads with high-access locality to further reduce restores. Third, we propose an adaptive restore mechanism to selectively pick the suitable restore scheme, according to the busy status of corresponding register banks. Experimental results show that our proposed design can effectively mitigate the performance loss and energy overhead caused by restore operations while still maintaining the reliability of reads.