High Performance MPI over the Slingshot Interconnect: Early Experiences

High Performance MPI over the Slingshot Interconnect: Early Experiences
复制标题

基于 Slingshot 互连的高性能 MPI:早期经验

DOI:
10.1145/3491418.3530773
复制
发表时间:
2022
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
Panda, Dhabaleswar
Panda, Dhabaleswar
中科院分区:
--
文献类型:
--
作者:
Shafie Khorassani, Kawthar;Chen, Chen Chun;Ramesh, Bharath;Shafi, Aamir;Subramoni, Hari;Panda, Dhabaleswar

文献摘要

参考文献

被引文献

相似文献

由HPE/Cray设计的Slingshot互连在即将推出的兆兆级系统上部署后,在高性能计算中的相关性越来越大。特别是,它是世界上第一台亿亿级和排名最高的超级计算机Frontier的互连。它提供各种功能,如自适应路由、拥塞控制和隔离工作负载。新互连的部署引发了有关性能、可扩展性和任何潜在瓶颈的问题,因为它们是这些系统上跨节点可扩展性的关键因素。在本文中,我们将深入研究弹弓互连与当前最先进的MPI库所带来的挑战。特别是,我们将研究跨节点使用slingshot时的可伸缩性性能。我们提出了一个全面的评估,使用各种MPI和通信库,包括Cray MPICH,OpenMPI + UCX,RCCL和MVAPICH 2-GDR的GPU上的Spock系统,抢先体验集群部署与弹弓和AMD MI 100 GPU,以模拟边疆系统。
The Slingshot interconnect designed by HPE/Cray is becoming more relevant in High-Performance Computing with its deployment on the upcoming exascale systems. In particular, it is the interconnect empowering the first exascale and highest-ranked supercomputer in the world, Frontier. It offers various features such as adaptive routing, congestion control, and isolated workloads. The deployment of newer interconnects raises questions about performance, scalability, and any potential bottlenecks as they are a critical element contributing to the scalability across nodes on these systems. In this paper, we will delve into the challenges the slingshot interconnect poses with current state-of-the-art MPI libraries. In particular, we look at the scalability performance when using slingshot across nodes. We present a comprehensive evaluation using various MPI and communication libraries including Cray MPICH, OpenMPI + UCX, RCCL, and MVAPICH2-GDR on GPUs on the Spock system, an early access cluster deployed with Slingshot and AMD MI100 GPUs, to emulate the Frontier system.
MVAPICH 项目:将研究转化为 HPC 社区的高性能 MPI 库
DOI: 10.1016/j.jocs.2020.101208
发表时间: 2020
影响因子: 3.3
作者:
Panda, Dhabaleswar Kumar;Subramoni, Hari;Chu, Ching-Hsiang;Bayatpour, Mohammadreza
通讯作者: Bayatpour, Mohammadreza
DOI: --
发表时间: 2020
期刊: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
D. D. Sensi;S. D. Girolamo;K. McMahon;D. Roweth;T. Hoefler
通讯作者: T. Hoefler
为 AMD GPU 设计 ROCm 感知 MPI 库:早期经验
DOI: 10.1007/978-3-030-78713-4_7
发表时间: 2021
期刊: International Conference on High Performance Computing 2021
影响因子: --
作者:
Shafie, K;Hashmi, J;Chu, C;Chen, C;Subramoni, H;Panda, D K.
通讯作者: Panda, D K.