Improving Thread-level Parallelism in GPUs Through Expanding Register File to Scratchpad Memory

Improving Thread-level Parallelism in GPUs Through Expanding Register File to Scratchpad Memory
复制标题

通过将寄存器文件扩展到暂存器内存来提高 GPU 中的线程级并行性

DOI:
10.1145/3280849
复制
发表时间:
2018-11
期刊:
Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Hailong Yang
Hailong Yang
中科院分区:
其他
文献类型:
--
作者:
Chao Yu;Yuebin Bai;Qingxiao Sun;Hailong Yang

文献摘要

参考文献

相似文献

现代图形处理器(GPU)由于其高性能和大规模线程级并行性(TLP)而成为计算机中的普适计算设备。GPU配备有大型寄存器文件(RF)以支持大规模线程之间的快速上下文切换,以及暂存器存储器(SPM)以支持协作线程阵列(CTA)内的线程间通信。然而,GPU的TLP通常受到寄存器文件和暂存器存储器的低效资源管理的限制。这种低效率还导致寄存器文件和暂存器存储器利用不足。为了克服上述效率低下,我们提出了一种新的资源管理方法EXPARS的GPU。EXPARS通过将寄存器文件扩展到暂存器存储器,在逻辑上提供了更大的寄存器文件。当可用的寄存器文件变得有限时,我们的方法利用未充分利用的暂存器来支持额外的寄存器分配。因此,可以将更多的CTA分派给SM,这提高了GPU利用率。我们在代表性基准测试套件上的实验表明,分配给每个SM的CTA数量平均增加了1.28倍。此外,我们的方法显着提高了GPU的资源利用率,寄存器文件利用率提高了11.64%,暂存器内存利用率平均提高了48.20%。有了更好的TLP,我们的方法平均实现了20.01%的性能提升,而能量开销可以忽略不计。
Modern Graphic Processing Units (GPUs) have become pervasive computing devices in datacenters due to their high performance with massive thread level parallelism (TLP). GPUs are equipped with large register files (RF) to support fast context switch between massive threads and scratchpad memory (SPM) to support inter-thread communication within the cooperative thread array (CTA). However, the TLP of GPUs is usually limited by the inefficient resource management of register file and scratchpad memory. This inefficiency also leads to register file and scratchpad memory underutilization. To overcome the above inefficiency, we propose a new resource management approach EXPARS for GPUs. EXPARS provides a larger register file logically by expanding the register file to scratchpad memory. When the available register file becomes limited, our approach leverages the underutilized scratchpad memory to support additional register allocation. Therefore, more CTAs can be dispatched to SMs, which improves the GPU utilization. Our experiments on representative benchmark suites show that the number of CTAs dispatched to each SM increases by 1.28× on average. In addition, our approach improves the GPU resource utilization significantly, with the register file utilization improved by 11.64% and the scratchpad memory utilization improved by 48.20% on average. With better TLP, our approach achieves 20.01% performance improvement on average with negligible energy overhead.
DOI: 10.1145/2967938.2967941
发表时间: 2016-09
期刊: 2016 International Conference on Parallel Architecture and Compilation Techniques (PACT)
影响因子: --
作者:
Onur Kayiran;Adwait Jog;Ashutosh Pattnaik;Rachata Ausavarungnirun;Xulong Tang;M. Kandemir;G. Loh;
通讯作者: Onur Kayiran;Adwait Jog;Ashutosh Pattnaik;Rachata Ausavarungnirun;Xulong Tang;M. Kandemir;G. Loh;
DOI: 10.1145/2628071.2628107
发表时间: 2014-08
期刊: 2014 23rd International Conference on Parallel Architecture and Compilation (PACT)
影响因子: --
作者:
Shin-Ying Lee;Carole-Jean Wu
通讯作者: Shin-Ying Lee;Carole-Jean Wu
DOI: 10.1145/3177964
发表时间: 2018-03
期刊: ACM Transactions on Architecture and Code Optimization (TACO)
影响因子: --
作者:
Zhen Lin;Mike Mantor;Huiyang Zhou
通讯作者: Zhen Lin;Mike Mantor;Huiyang Zhou
DOI: 10.1145/2907294.2907298
发表时间: 2015-03
期刊: Proceedings of the 25th ACM International Symposium on High-Performance Parallel and Distributed Computing
影响因子: --
作者:
Vishwesh Jatala;Jayvant Anantpur;Amey Karkare
通讯作者: Vishwesh Jatala;Jayvant Anantpur;Amey Karkare
DOI: 10.1109/4.509850
发表时间: 1996-05-01
影响因子: 5.4
作者:
Wilton, SJE;Jouppi, NP
通讯作者: Jouppi, NP