Exploiting Maximal Overlap for Non-Contiguous Data Movement Processing on Modern GPU-Enabled Systems

Exploiting Maximal Overlap for Non-Contiguous Data Movement Processing on Modern GPU-Enabled Systems
复制标题

在现代 GPU 系统上利用最大重叠进行非连续数据移动处理

DOI:
10.1109/ipdps.2016.99
复制
发表时间:
2016
期刊:
2016 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Ching;Khaled Hamidouche;Akshay Venkatesh;D. Banerjee;H. Subramoni;D. Panda

文献摘要

被引文献

相似文献

GPU加速器由于其大规模并行性和高每瓦吞吐量而广泛用于HPC集群。数据移动仍然是GPU集群的主要瓶颈,当数据不连续时更是如此,这在科学应用中很常见。CUDA-Aware MPI库使用面向延迟的技术优化非连续数据移动处理,例如使用GPU内核来加速打包/解包操作。虽然它们优化了单个操作的延迟,但是设计的固有限制限制了它们对于面向吞吐量的模式的效率。事实上,现有的设计都没有充分利用GPU的大规模并行性,以通过实现最大重叠来提供高吞吐量和有效的资源利用率。在本文中,我们提出了新的CUDA感知MPI库设计,以实现高效的GPU资源利用率和CPU与GPU之间的最大重叠,用于非连续数据处理和移动。所提出的设计利用了几个CUDA功能,如Hyper-Q/多流和回调函数,以提供高性能和效率。据我们所知,这是第一项为非连续MPI数据处理和往返GPU的移动提供高吞吐量和高效资源利用的此类研究。使用DDTBench对所提出的设计进行的性能评估显示,对于节点内GPU间乒乓实验,SPECFEM3D_oc,SPECFEM3D_cm和WRF_y_sa的性能分别提高了54%,67%和61%。所提出的设计还提供了高达33%的总执行时间比现有的设计HaloExchange为基础的应用程序内核,模型的通信模式的MeteoSwiss天气预报模型超过32个GPU节点上的威尔克斯GPU集群的改进。
GPU accelerators are widely used in HPC clusters due to their massive parallelism and high throughput-per-watt. Data movement continues to be the major bottleneck on GPU clusters, more so when data is non-contiguous, which is common in scientific applications. CUDA-Aware MPI libraries optimize the non-contiguous data movement processing using latency oriented techniques such as using GPU kernels to accelerate the packing/unpacking operations. Although they optimize the latency of a single operation, the inherent restrictions of the designs limit their efficiency for throughput oriented patterns. Indeed, none of the existing designs fully exploit the massive parallelism of the GPUs to provide high throughput and efficient resources utilization by enabling maximal overlap. In this paper, we propose novel designs for CUDA-Aware MPI libraries to achieve efficient GPU resource utilization and maximal overlap between CPUs and GPUs for non-contiguous data processing and movement. The proposed designs take advantage of several CUDA features, such as Hyper-Q/multi-streams and callback function, to deliver high performance and efficiency. To the best of our knowledge, this is the first such study to provide high throughput and efficient resource utilization for non-contiguous MPI data processing and movement to/from GPUs. The performance evaluation with the proposed designs using DDTBench shows up to 54%, 67%, 61% performance improvement on the SPECFEM3D_oc, SPECFEM3D_cm and WRF_y_sa benchmarks respectively for intra-node inter-GPU ping-pong experiments. The proposed designs also deliver up to 33% improvement on the total execution time over the existing designs for the HaloExchange-based application kernel that models the communication pattern of the MeteoSwiss weather forecasting model over 32 GPU nodes on Wilkes GPU cluster.