Automatic Restructuring of GPU Kernels for Exploiting Inter-thread Data Locality
Automatic Restructuring of GPU Kernels for Exploiting Inter-thread Data Locality
复制标题
自动重构 GPU 内核以利用线程间数据局部性
DOI:
10.1007/978-3-642-28652-0_2
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Apan Qasem
中科院分区:
文献类型:
--
作者:
Swapneela Unkule;Christopher Shaltz;Apan Qasem
Hundreds of cores per chip and support for fine-grain multithreading have made GPUs a central player in today's HPC world. For many applications, however, achieving a high fraction of peak on current GPUs, still requires significant programmer effort. A key consideration for optimizing GPU code is determining a suitable amount of work to be performed by each thread. Thread granularity not only has a direct impact on occupancy but can also influence data locality at the register and shared-memory levels. This paper describes a software framework to analyze dependencies in parallel GPU threads and perform source-level restructuring to obtain GPU kernels with varying thread granularity. The framework supports specification of coarsening factors through source-code annotation and also implements a heuristic based on estimated register pressure that automatically recommends coarsening factors for improved memory performance. We present preliminary experimental results on a select set of CUDA kernels. The results show that the proposed strategy is generally able to select profitable coarsening factors. More importantly, the results demonstrate a clear need for automatic control of thread granularity at the software level for achieving higher performance.