Enabling Fast, Noncontiguous GPU Data Movement in Hybrid MPI+GPU Environments

Enabling Fast, Noncontiguous GPU Data Movement in Hybrid MPI+GPU Environments
复制标题

DOI:
10.1109/cluster.2012.72
复制
发表时间:
2012-09
期刊:
2012 IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
John Jenkins;James Dinan;P. Balaji;N. Samatova;R. Thakur
John Jenkins;James Dinan;P. Balaji;N. Samatova;R. Thakur
中科院分区:
其他
文献类型:
--
作者:
John Jenkins;James Dinan;P. Balaji;N. Samatova;R. Thakur

文献摘要

被引文献

相似文献

在MPI + GPU混合环境中缺乏与GPU数据的高效透明交互,这对大规模科学计算的GPU加速提出了挑战。一个特殊的挑战是将非连续数据传输到GPU内存和从GPU内存传输。MPI实现当前不提供利用数据类型用于GPU存储器中的数据的非连续通信的有效手段。为了解决这个问题,我们提出了一个MPI数据类型处理系统,能够有效地处理任意数据类型直接在GPU上。我们提出了一种将传统的数据类型表示转换为GPU兼容的格式。然后,GPU内核利用细粒度的元素级并行性来执行非连续元素的设备内打包和解包。我们证明了几倍的性能改进,非连续列向量,3D阵列切片,4D阵列子卷基于CUDA的替代品。与优化的,布局特定的实现相比,我们的方法产生的开销低,同时使包装的数据类型,没有一个直接的CUDA等效。这些改进被证明可以转化为端到端、GPU到GPU通信时间的显著改进。此外,我们确定和评估的通信模式,可能会导致资源竞争与包装操作,提供了一个基线,自适应地选择数据处理策略。
Lack of efficient and transparent interaction with GPU data in hybrid MPI+GPU environments challenges GPU acceleration of large-scale scientific computations. A particular challenge is the transfer of noncontiguous data to and from GPU memory. MPI implementations currently do not provide an efficient means of utilizing data types for noncontiguous communication of data in GPU memory. To address this gap, we present an MPI data type-processing system capable of efficiently processing arbitrary data types directly on the GPU. We present a means for converting conventional data type representations into a GPU-amenable format. Fine-grained, element-level parallelism is then utilized by a GPU kernel to perform in-device packing and unpacking of noncontiguous elements. We demonstrate a several-fold performance improvement for noncontiguous column vectors, 3D array slices, and 4D array sub volumes over CUDA-based alternatives. Compared with optimized, layout-specific implementations, our approach incurs low overhead, while enabling the packing of data types that do not have a direct CUDA equivalent. These improvements are demonstrated to translate to significant improvements in end-to-end, GPU-to-GPU communication time. In addition, we identify and evaluate communication patterns that may cause resource contention with packing operations, providing a baseline for adaptively selecting data-processing strategies.