Improving MPI Multi-threaded RMA Communication Performance

Improving MPI Multi-threaded RMA Communication Performance
复制标题

提高MPI多线程RMA通信性能

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Parallel Processing
影响因子:
--
通讯作者:
D. Arnold
D. Arnold
中科院分区:
--
文献类型:
--
作者:
N. Hjelm;Matthew G. F. Dosanjh;Ryan E. Grant;Taylor L. Groves;P. Bridges;D. Arnold

文献摘要

被引文献

相似文献

单边通信对于实现通信并发性至关重要。随着内核数量的增加,特别是在众核架构中,单边(RMA)通信被提出来解决网络接口处不断增加的争用。在MPI中使用单边(RMA)通信的困难在于,使用RMA和多个并发线程的MPI实现的性能还没有得到很好的理解。过去的研究已经使用MPI RMA结合多线程(RMA-MT),但他们已经在旧的MPI实现缺乏RMA-MT优化。此外,先前的工作仅在较小规模(<=512个核心)下进行。在本文中,我们描述了一个新的RMA实现开放MPI。实现的目标是可伸缩性和多线程性能。我们描述了我们的RMA改进的设计和实施,并提供了一个评估,演示了扩展到524,288个核心,领先的超级计算机安装的全尺寸。相比之下,之前的实现未能扩展到超过大约4,096个核心。为了评估这种方法,我们比较了供应商优化的MPI RMA-MT实现与微基准,一个迷你应用程序,并在大规模的多核架构上的完整天体物理代码。这是第一次对MPI RMA-MT(524,288个核心)的众核架构进行大规模评估,也是第一次在两种不同的RMA-MT优化MPI实现之间进行大规模应用性能比较。结果显示,对于在512 K内核上运行的完整应用程序代码,我们优化的开源MPI的收益为8.6%。
One-sided communication is crucial to enabling communication concurrency. As core counts have increased, particularly with many-core architectures, one-sided (RMA) communication has been proposed to address the ever increasing contention at the network interface. The difficulty in using one-sided (RMA) communication with MPI is that the performance of MPI implementations using RMA with multiple concurrent threads is not well understood. Past studies have been done using MPI RMA in combination with multi-threading (RMA-MT) but they have been performed on older MPI implementations lacking RMA-MT optimizations. In addition prior work has only been done at smaller scale (<=512 cores). In this paper, we describe a new RMA implementation for Open MPI. The implementation targets scalability and multi-threaded performance. We describe the design and implementation of our RMA improvements and offer an evaluation that demonstrates scaling to 524,288 cores, the full size of a leading supercomputer installation. In contrast, the previous implementation failed to scale past approximately 4,096 cores. To evaluate this approach, we then compare against a vendor optimized MPI RMA-MT implementation with microbenchmarks, a mini-application, and a full astrophysics code at large scale on a many-core architecture. This is the first time that an evaluation at large scale on many-core architectures has been done for MPI RMA-MT (524,288 cores) and the first large scale application performance comparison between two different RMA-MT optimized MPI implementations. The results show a 8.6% benefit to our optimized open source MPI for a full application code running on 512K cores.