Automatic Restructuring of GPU Kernels for Exploiting Inter-thread Data Locality

Automatic Restructuring of GPU Kernels for Exploiting Inter-thread Data Locality
复制标题

自动重构 GPU 内核以利用线程间数据局部性

DOI:
10.1007/978-3-642-28652-0_2
复制
发表时间:
2012
期刊:
2021 IEEE/ACM International Symposium on Code Generation and Optimization (CGO)
影响因子:
--
通讯作者:
Apan Qasem
Apan Qasem
中科院分区:
--
文献类型:
--
作者:
Swapneela Unkule;Christopher Shaltz;Apan Qasem

文献摘要

被引文献

相似文献

每个芯片的数百个内核和对细粒度多线程的支持使gpu成为当今高性能计算领域的核心玩家。然而,对于许多应用程序来说,在当前gpu上实现高峰值仍然需要大量程序员的努力。优化GPU代码的一个关键考虑因素是确定每个线程要执行的适当工作量。线程粒度不仅对占用有直接影响,而且还会影响寄存器和共享内存级别的数据位置。本文描述了一个软件框架,用于分析并行GPU线程中的依赖关系,并执行源代码级重构以获得不同线程粒度的GPU内核。该框架支持通过源代码注释规范粗化因子,并且还实现了基于估计寄存器压力的启发式方法,该方法自动推荐粗化因子以提高内存性能。我们在一组选定的CUDA内核上给出了初步实验结果。结果表明,该策略总体上能够选择出有利的粗化因子。更重要的是,结果清楚地表明,为了实现更高的性能,需要在软件级别自动控制线程粒度。
Hundreds of cores per chip and support for fine-grain multithreading have made GPUs a central player in today's HPC world. For many applications, however, achieving a high fraction of peak on current GPUs, still requires significant programmer effort. A key consideration for optimizing GPU code is determining a suitable amount of work to be performed by each thread. Thread granularity not only has a direct impact on occupancy but can also influence data locality at the register and shared-memory levels. This paper describes a software framework to analyze dependencies in parallel GPU threads and perform source-level restructuring to obtain GPU kernels with varying thread granularity. The framework supports specification of coarsening factors through source-code annotation and also implements a heuristic based on estimated register pressure that automatically recommends coarsening factors for improved memory performance. We present preliminary experimental results on a select set of CUDA kernels. The results show that the proposed strategy is generally able to select profitable coarsening factors. More importantly, the results demonstrate a clear need for automatic control of thread granularity at the software level for achieving higher performance.