Adaptive and Hierarchical Large Message All-to-all Communication Algorithms for Large-scale Dense GPU Systems

Adaptive and Hierarchical Large Message All-to-all Communication Algorithms for Large-scale Dense GPU Systems
复制标题

DOI:
10.1109/ccgrid51090.2021.00021
复制
发表时间:
2021-05
期刊:
2021 IEEE/ACM 21st International Symposium on Cluster, Cloud and Internet Computing (CCGrid)
影响因子:
--
通讯作者:
Kawthar Shafie Khorassani;Ching-Hsiang Chu;Quentin G. Anthony;H. Subramoni;D. Panda
Kawthar Shafie Khorassani;Ching-Hsiang Chu;Quentin G. Anthony;H. Subramoni;D. Panda
中科院分区:
其他
文献类型:
--
作者:
Kawthar Shafie Khorassani;Ching-Hsiang Chu;Quentin G. Anthony;H. Subramoni;D. Panda

文献摘要

被引文献

相似文献

近年来,GPU增强的集群在高性能计算(HPC)中变得越来越普遍,导致对更有效的多GPU通信的需求。这使得探索可以通过MPI等通信中间件获得的性能增强越来越重要,以充分利用这些系统上可用的GPU。在本文中,我们提出了在大规模密集的GPU系统上全面的全部集体沟通的地方感知和自适应方案。所提出的算法利用了通过GPU之间的NVLink互连提供的高带宽,以重叠通信延迟。我们专注于个性化和非人性化的全体集体沟通。这些是现代科学计算应用的组成部分,利用矩阵转台和三维快速傅立叶变换(FFT),并且与模型和混合平行性的深度学习工作负载变得更加相关。执行三维FFT的应用程序内核的性能评估表明,全能的个性化方案可能会导致Lassen系统上256 GPU的执行时间降低15-25%。我们证明了在深度学习培训中使用的分布式K-FAC的训练时间大约提高了约8%,该培训最多可达128 GPU。与最先进的MPI库和Lassen Systems相比,我们还分别证明了非个人化和个性化的全能基准的性能分别提高了大约22%和30%。
In recent years, GPU-enhanced clusters have become more prevalent in High-Performance Computing (HPC), leading to a demand for more efficient multi-GPU communication. This makes it increasingly important to explore performance enhancements that can be attained through the communication middleware such as MPI, in order to fully take advantage of the GPUs available on these systems. In this paper, we propose locality-aware and adaptive schemes for hierarchical All-to-all collective communication on large-scale dense GPU systems. The proposed algorithms utilize the high bandwidth made available through the NVLink interconnect between GPUs in order to overlap communication latency. We focus on personalized and non-personalized all-to-all collective communication. These are components of modern scientific computing applications that utilize matrix transpose and three-dimensional Fast Fourier Transforms (FFT) and becoming more relevant for Deep Learning workloads with model and hybrid parallelisms. The performance evaluation with an application kernel performing three-dimensional FFT indicates that the proposed schemes for personalized all-to-all can lead to up to 15-25% lower execution time on 256 GPUs on the Lassen system. We demonstrate approximately 8% enhancement in training time for distributed K-FAC used in Deep Learning training on up to 128 GPUs. We also demonstrate approximately 22% and 30% improvement in the performance of non-personalized and personalized all-to-all benchmarks, respectively, compared to the state-of-the-art MPI libraries on the Summit and Lassen systems.