GPU-aware Communication with UCX in Parallel Programming Models: Charm++, MPI, and Python

GPU-aware Communication with UCX in Parallel Programming Models: Charm++, MPI, and Python
复制标题

DOI:
10.1109/ipdpsw52791.2021.00079
复制
发表时间:
2021-06
期刊:
2021 IEEE International Parallel and Distributed Processing Symposium Workshops (IPDPSW)
影响因子:
--
通讯作者:
Jaemin Choi;Zane Fink;Sam White;Nitin Bhat;D. Richards;L. Kalé
Jaemin Choi;Zane Fink;Sam White;Nitin Bhat;D. Richards;L. Kalé
中科院分区:
其他
文献类型:
--
作者:
Jaemin Choi;Zane Fink;Sam White;Nitin Bhat;D. Richards;L. Kalé

文献摘要

被引文献

相似文献

随着越来越多的领导级系统采用GPU加速器,GPU数据的高效通信正在成为高性能计算最关键的组成部分之一。对于并行编程模型的开发人员来说,使用CUDA等gpu的本机api实现对gpu感知通信的支持可能是一项艰巨的任务,因为它需要相当大的努力,几乎不能保证性能。在这项工作中,我们展示了统一通信X (UCX)框架组成gpu感知通信层的能力,该通信层服务于Charm++生态系统的多个并行编程模型:Charm++,自适应MPI (AMPI)和Charm4py。我们使用从OSU基准套件改编的微基准测试来演示我们的设计对性能的影响,在Charm++、AMPI和Charm4py中分别获得了高达10.2倍、11.7倍和17.4倍的延迟改进。我们还观察到,在Charm++中带宽增加了9.6倍,在AMPI中增加了10倍,在Charm4py中增加了10.5倍。通过评估Jacobi迭代方法的代理应用程序,我们展示了我们的设计对实际应用程序的潜在影响,在Charm++中提高了12.4倍的通信性能,在AMPI中提高了12.8倍,在Charm4py中提高了19.7倍。
As an increasing number of leadership-class systems embrace GPU accelerators in the race towards exascale, efficient communication of GPU data is becoming one of the most critical components of high-performance computing. For developers of parallel programming models, implementing support for GPU-aware communication using native APIs for GPUs such as CUDA can be a daunting task as it requires considerable effort with little guarantee of performance. In this work, we demonstrate the capability of the Unified Communication X (UCX) framework to compose a GPU-aware communication layer that serves multiple parallel programming models of the Charm++ ecosystem: Charm++, Adaptive MPI (AMPI), and Charm4py. We demonstrate the performance impact of our designs with microbenchmarks adapted from the OSU benchmark suite, obtaining improvements in latency of up to 10.2x, 11.7x, and 17.4x in Charm++, AMPI, and Charm4py, respectively. We also observe increases in bandwidth of up to 9.6x in Charm++, 10x in AMPI, and 10.5x in Charm4py. We show the potential impact of our designs on real-world applications by evaluating a proxy application for the Jacobi iterative method, improving the communication performance by up to 12.4x in Charm++, 12.8x in AMPI, and 19.7x in Charm4py.