Countering Load-to-Use Stalls in the NVIDIA Turing GPU

Countering Load-to-Use Stalls in the NVIDIA Turing GPU
复制标题

解决 NVIDIA Turing GPU 中的加载使用停顿问题

DOI:
10.1109/mm.2020.3012514
复制
发表时间:
2020
期刊:
影响因子:
3.6
通讯作者:
Alexandre Joly
Alexandre Joly
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ram Rangan;Naman Turakhia;Alexandre Joly

文献摘要

被引文献

相似文献

在其对先前NVIDIA GPU的各种改进中,NVIDIA Turing GPU具有四个关键性能增强功能,可有效地反对内存负载对使用的失速。首先,降低了全局内存负载的L1命中的延迟有助于较低的平均内存查找延迟。接下来,能够在可缓存的内存和SCRATCHPAD或共享内存之间动态配置L1数据RAM的能力,使驱动程序软件能够为具有低共享内存需求,增加L1命中并减少负载到使用的载荷停顿的程序的L1数据缓存大小最大化L1 Data Cache大小。最后,矢量寄存器文件容量加倍的双重增强功能以​​及添加专用的标量或统一寄存器文件以及统一数据,均匀的数据,轻松矢量寄存器压力并启用更高的扭曲级并行性,从而提高延迟公差。我们发现,上述增强功能总结为现代游戏应用程序的平均速度为11%。
Among its various improvements over prior NVIDIA GPUs, the NVIDIA Turing GPU boasts of four key performance enhancements to effectively counter memory load-to-use stalls. First, reduced latency on L1 hits for global memory loads helps lower average memory lookup latency. Next, the ability to dynamically configure the L1 data RAM between cacheable memory and scratchpad or shared memory, enables driver software to opportunistically maximize L1 data cache size for programs with low shared memory requirements, increasing L1 hits and reducing load-to-use stalls. Finally, the twin enhancements of doubling of vector register file capacity and the addition of a dedicated scalar or uniform register file along with a uniform datapath, ease vector register pressure and enable higher warp level parallelism, leading to better latency tolerance. We find that the above enhancements combined deliver an average speedup of 11% on modern gaming applications.