C-GDR: High-Performance Container-Aware GPUDirect MPI Communication Schemes on RDMA Networks

C-GDR: High-Performance Container-Aware GPUDirect MPI Communication Schemes on RDMA Networks
复制标题

DOI:
10.1109/ipdps.2019.00034
复制
发表时间:
2019-05
期刊:
2019 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Jie Zhang;Xiaoyi Lu;Ching-Hsiang Chu;D. Panda
Jie Zhang;Xiaoyi Lu;Ching-Hsiang Chu;D. Panda
中科院分区:
其他
文献类型:
--
作者:
Jie Zhang;Xiaoyi Lu;Ching-Hsiang Chu;D. Panda

文献摘要

被引文献

相似文献

近年来,基于 GPU 的平台在并行应用程序方面取得了巨大成功。除了 GPU 上高度优化的计算内核之外,GPU 集群上的数据移动成本在为最终应用程序提供高性能方面也发挥着关键作用。最近提出了许多研究来优化 GPU 或 CUDA 感知通信运行时的性能,并且这些设计已广泛应用于新兴的基于 GPU 的应用程序中。这些研究主要集中在提高本地环境(即物理机)上的通信性能,但是云环境上基于 GPU 的通信方案尚未得到很好的研究。本文首先研究了最先进的基于 GPU 的通信方案在本机和基于容器的环境中的性能特征,这表明在支持 GPU 的运行时中设计高性能容器感知通信方案的巨大需求,以便为云上的终端应用程序提供接近本机的性能。接下来,我们提出 C-GDR 方法来设计 RDMA 网络上的高性能容器感知 GPUDirect 通信方案。 C-GDR 允许通信运行时成功检测进程局部性、GPU 驻留、NUMA、架构信息和通信模式,以便在支持 GPU 的云上智能、动态地选择最佳通信和数据移动方案。我们已将 C-GDR 与 MAPICH2 库集成。我们的评估表明,与默认的 MVAPICH2-GDR 和 Open MPI 相比,具有 C-GDR 的 MVAPICH2 在基于容器的云环境中具有明显的性能优势。例如,我们提出的 C-GDR 在微基准测试中比默认的 MVAPICH2-GDR 方案性能高出 66%,在基于容器的环境中比 HPC 应用程序高出 26%。
In recent years, GPU-based platforms have received significant success for parallel applications. In addition to highly optimized computation kernels on GPUs, the cost of data movement on GPU clusters plays critical roles in delivering high performance for end applications. Many recent studies have been proposed to optimize the performance of GPU-or CUDA-aware communication runtimes and these designs have been widely adopted in the emerging GPU-based applications. These studies mainly focus on improving the communication performance on native environments, i.e., physical machines, however GPU-based communication schemes on cloud environments are not well studied yet. This paper first investigates the performance characteristics of state-of-the-art GPU-based communication schemes on both native and container-based environments, which show a significant demand to design high-performance container-aware communication schemes in GPU-enabled runtimes to deliver near-native performance for end applications on clouds. Next, we propose the C-GDR approach to design high-performance Container-aware GPUDirect communication schemes on RDMA networks. C-GDR allows communication runtimes to successfully detect process locality, GPU residency, NUMA, architecture information, and communication pattern to enable intelligent and dynamic selection of the best communication and data movement schemes on GPU-enabled clouds. We have integrated C-GDR with the MVAPICH2 library. Our evaluations show that MVAPICH2 with C-GDR has clear performance benefits on container-based cloud environments, compared to default MVAPICH2-GDR and Open MPI. For instance, our proposed C-GDR can outperform default MVAPICH2-GDR schemes by up to 66% on micro-benchmarks and up to 26% on HPC applications over a container-based environment.