Performance Evaluation of MPI Libraries on GPU-Enabled OpenPOWER Architectures: Early Experiences

Performance Evaluation of MPI Libraries on GPU-Enabled OpenPOWER Architectures: Early Experiences
复制标题

支持 GPU 的 OpenPOWER 架构上 MPI 库的性能评估:早期经验

DOI:
10.1007/978-3-030-34356-9_28
复制
发表时间:
2019
期刊:
The American journal of medicine
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Kawthar Shafie Khorassani;Ching;H. Subramoni;D. Panda

文献摘要

被引文献

相似文献

支持图形处理单元(GPU)的OpenPOWER架构的出现,推动了各种高性能计算(HPC)应用的发展,从动态模块化模拟到深度学习训练。gpu感知的消息传递接口(MPI)是用于大规模利用支持gpu的HPC系统上的计算能力的最有效的库之一。然而,缺乏对支持gpu的MPI库进行全面的性能评估,以深入了解在支持gpu的OpenPOWER系统上使用每个MPI库的不同成本和收益。在本文中,我们提供了详细的性能评估和点对点通信的分析,使用各种gpu感知的MPI库,包括SpectrumMPI, OpenMPI+UCX和MVAPICH2-GDR在OpenPOWER支持gpu的系统上。我们证明,所有三个MPI库为同一套接字上的两个gpu之间的NVLink通信提供了大约95%的可实现带宽。对于InfiniBand网络主导峰值带宽的节点间通信,MVAPICH2-GDR和SpectrumMPI实现了大约99%的可实现带宽,而OpenMPI提供了接近95%的可实现带宽。此评估对于确定哪个MPI库可以提供最高的性能增强非常有用。
The advent of Graphics Processing Unit (GPU)-enabled OpenPOWER architectures are empowering the advancement of various High-Performance Computing (HPC) applications from dynamic modular simulation to deep learning training. GPU-aware Message Passing Interface (MPI) is one of the most efficient libraries used to exploit the computing power on GPU-enabled HPC systems at scale. However, there is a lack of thorough performance evaluations for GPU-aware MPI libraries to provide insights into the varying costs and benefits of using each one on GPU-enabled OpenPOWER systems. In this paper, we provide a detailed performance evaluation and analysis of point-to-point communication using various GPU-aware MPI libraries including SpectrumMPI, OpenMPI+UCX, and MVAPICH2-GDR on OpenPOWER GPU-enabled systems. We demonstrate that all three MPI libraries deliver approximately 95% of achievable bandwidth for NVLink communication between two GPUs on the same socket. For inter-node communication where the InfiniBand network dominates the peak bandwidth, MVAPICH2-GDR and SpectrumMPI attain approximately 99% achievable bandwidth, while OpenMPI delivers close to 95%. This evaluation is useful to determine which MPI library can provide the highest performance enhancement.