Efficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters

Efficient Personalized and Non-Personalized Alltoall Communication for Modern Multi-HCA GPU-Based Clusters
复制标题

DOI:
10.1109/hipc56025.2022.00025
复制
发表时间:
2022-12
期刊:
2022 IEEE 29th International Conference on High Performance Computing, Data, and Analytics (HiPC)
影响因子:
--
通讯作者:
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda
中科院分区:
其他
文献类型:
--
作者:
K. Suresh;Akshay Paniraja Guptha;Benjamin Michalowicz;B. Ramesh;M. Abduljabbar;A. Shafi;H. Subramoni;D. Panda

文献摘要

相似文献

图形处理单元(GPU)在当今的超级计算集群中已变得无处不在基于GPU的系统以前具有多个HCA,科学家利用多HCA系统来加速CPU之间的节点转移在这项工作中,我们需要使用MPI_Allgather作为示例,我们建议使用MPI_ALLGATHER,我们提出了一个有效的MPI_Allgather算法,并将其扩展到MPI_AlltoAll使用OMB基准套件的算法的性能分别在非个人化和个性化的通信基准中提高30%和43%与128 GPU的最先进的MPI库相比
Graphics Processing Units (GPUs) have become ubiquitous in today’s supercomputing clusters primarily because of their high compute capability and power efficiency. Message Passing Interface (MPI) is a widely adopted programming model for large-scale GPU-based applications used in such clusters. Modern GPU-based systems have multiple HCAs. Previously, scientists have leveraged multi-HCA systems to accelerate inter-node transfers between CPUs using point-to-point primitives. In this work, we show the need for collective-level, multi-rail aware algorithms using MPI_Allgather as an example. We then propose an efficient multi-rail MPI_Allgather algorithm and extend it to MPI_Alltoall. We analyze the performance of this algorithm using OMB benchmark suite. We demonstrate approximately 30% and 43% improvement in non-personalized and personalized communication benchmarks respectively when compared with the state-of-the-art MPI libraries on 128 GPUs