Efficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters
Efficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters
复制标题
DOI:
10.1109/hipc56025.2022.00025
复制
发表时间:
2022-12
期刊:
影响因子:
--
通讯作者:
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda
中科院分区:
文献类型:
--
作者:
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda
Graphics Processing Units (GPUs) have become ubiquitous in today’s supercomputing clusters primarily because of their high compute capability and power efficiency. Message Passing Interface (MPI) is a widely adopted programming model for large-scale GPU-based applications used in such clusters. Modern GPU-based systems have multiple HCAs. Previously, scientists have leveraged multi-HCA systems to accelerate inter-node transfers between CPUs using point-to-point primitives. In this work, we show the need for collective-level, multi-rail aware algorithms using MPI_Allgather as an example. We then propose an efficient multi-rail MPI_Allgather algorithm and extend it to MPI_Alltoall. We analyze the performance of this algorithm using OMB benchmark suite. We demonstrate approximately 30% and 43% improvement in non-personalized and personalized communication benchmarks respectively when compared with the state-of-the-art MPI libraries on 128 GPUs