Data layout optimization for GPGPU architectures

Data layout optimization for GPGPU architectures
复制标题

GPGPU架构的数据布局优化

DOI:
--
复制
发表时间:
2013
期刊:
ACM SIGPLAN Symposium on Principles & Practice of Parallel Programming
影响因子:
--
通讯作者:
M. Kandemir
M. Kandemir
中科院分区:
--
文献类型:
--
作者:
Jun Liu;W. Ding;Ohyoung Jang;M. Kandemir

文献摘要

被引文献

相似文献

GPU被广泛用于加速通用应用程序,导致GPGPU架构的出现。新的编程模型,例如,计算统一设备架构(CUDA)已经被提出来促进在GPGPU中编程通用计算。然而,手动编写高性能CUDA代码仍然是繁琐和困难的。特别地,由于定制GPGPU存储器层次结构的独特特征,存储器空间中的数据的组织可以极大地影响性能。在这项工作中,我们提出了一个自动数据布局转换框架,以解决与GPGPU内存层次结构相关的关键问题(即,信道偏斜、数据合并和库冲突)。我们的方法采用了一种广泛适用的策略,基于一个新的概念,称为数据本地化。具体来说,我们试图优化布局的仿射循环嵌套访问的数组,设备内存和共享内存,在粗粒度和细粒度并行化水平。我们在NVIDIA CUDA GPU设备上使用15个基准测试对我们的数据布局优化策略进行了实验评估。结果表明,所提出的数据转换方法平均带来了4.3倍的加速比。
GPUs are being widely used in accelerating general-purpose applications, leading to the emergence of GPGPU architectures. New programming models, e.g., Compute Unified Device Architecture (CUDA), have been proposed to facilitate programming general-purpose computations in GPGPUs. However, writing high-performance CUDA codes manually is still tedious and difficult. In particular, the organization of the data in the memory space can greatly affect the performance due to the unique features of a custom GPGPU memory hierarchy. In this work, we propose an automatic data layout transformation framework to solve the key issues associated with a GPGPU memory hierarchy (i.e., channel skewing, data coalescing, and bank conflicts). Our approach employs a widely applicable strategy based on a novel concept called data localization. Specifically, we try to optimize the layout of the arrays accessed in affine loop nests, for both the device memory and shared memory, at both coarse grain and fine grain parallelization levels. We performed an experimental evaluation of our data layout optimization strategy using 15 benchmarks on an NVIDIA CUDA GPU device. The results show that the proposed data transformation approach brings around 4.3X speedup on average.