Enabling highly-scalable remote memory access programming with MPI-3 one sided

Enabling highly-scalable remote memory access programming with MPI-3 one sided
复制标题

使用 MPI-3 一侧实现高度可扩展的远程内存访问编程

DOI:
10.1145/2503210.2503286
复制
发表时间:
2013
期刊:
2013 SC - International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
Torsten Hoefler
Torsten Hoefler
中科院分区:
--
文献类型:
--
作者:
Robert Gerstenberger;Maciej Besta;Torsten Hoefler

文献摘要

被引文献

相似文献

现代互连提供远程直接内存访问(RDMA)功能。然而,大多数应用程序依赖于显式消息传递进行通信,尽管它们有不必要的开销。MPI-3.0标准定义了一个直接利用RDMA网络的编程接口,但其可扩展性和实用性需要在实践中得到验证。在这项工作中,我们开发了实现MPI-3.0规范的可扩展无缓冲协议。我们的协议支持扩展到数百万个内核,内存消耗可以忽略不计,同时提供最高的性能和最小的开销。为了武装程序员,我们为所有关键功能提供了一系列性能模型,并通过几个多达50万个进程的应用程序研究演示了我们的库和模型的可用性。我们表明,我们的设计在延迟、带宽和消息速率方面与UPC和Fortran阵列相当,甚至更好。我们还演示了具有相当编程复杂性的应用程序性能改进。
Modern interconnects offer remote direct memory access (RDMA) features. Yet, most applications rely on explicit message passing for communications albeit their unwanted overheads. The MPI-3.0 standard defines a programming interface for exploiting RDMA networks directly, however, it's scalability and practicability has to be demonstrated in practice. In this work, we develop scalable bufferless protocols that implement the MPI-3.0 specification. Our protocols support scaling to millions of cores with negligible memory consumption while providing highest performance and minimal overheads. To arm programmers, we provide a spectrum of performance models for all critical functions and demonstrate the usability of our library and models with several application studies with up to half a million processes. We show that our design is comparable to, or better than UPC and Fortran Coarrays in terms of latency, bandwidth, and message rate. We also demonstrate application performance improvements with comparable programming complexity.