High Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits

High Performance MPI Datatype Support with User-Mode Memory Registration: Challenges, Designs, and Benefits
复制标题

具有用户模式内存注册的高性能 MPI 数据类型支持:挑战、设计和优势

DOI:
--
复制
发表时间:
2015
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Mingzhe Li;H. Subramoni;Khaled Hamidouche;Xiaoyi Lu;D. Panda

文献摘要

被引文献

相似文献

非连续数据通信在科学应用中被大量采用,特别是对于那些用MPI编写的应用程序。处理非连续数据的常见策略,如打包/解包,在通信期间会产生显著的性能开销,这可能成为使用MPI派生数据类型的障碍。最近,Mellanox InfiniBand引入了一个新功能,称为用户模式内存注册(UMR),用于非连续数据通信。UMR有可能有效地支持MPI派生的数据类型通信,而没有打包/解包的开销。本文分析了UMR的特征,并利用InfiniBand动词级微基准研究了UMR的基本性能。在此基础上,我们提出了基于UMR的方案来支持MPI级的零拷贝数据类型通信。我们表明,UMR与MPI堆栈的天真集成不能带来比现有方案更好的性能优势。因此,我们提出了两种方案--UMR池和UMR缓存--以实现与UMR的高性能MPI数据类型通信。据我们所知,这是第一篇研究、分析和设计使用UMR功能的MPI非连续数据通信的文章。我们在MVAPICH2库的基础上提出并实现了基于UMR的设计。在微基准测试级别上的实验结果表明,与打包/解包方案相比,提出的基于UMR的设计能够将大消息向量基准测试的时延提高4倍。在应用程序级,对于在512个进程上具有MPI派生数据类型的3D模板通信内核,基于UMR的优化设计的执行时间比打包/解包方案的执行时间快27%。
Noncontiguous data communication has been heavily adopted in scientific applications, especially for those written with MPI. Common strategies to handle noncontiguous data, like packing/unpacking, incur significant performance overhead during communication, which could become as a barrier of using MPI derived datatypes. Recently, a novel feature of Mellanox InfiniBand, called User-mode Memory Registration (UMR), has been introduced for noncontiguous data communication. UMR has the potential to support MPI derived datatype communication efficiently without the overhead of packing/unpacking. In this paper, we analyze the UMR feature and study its basic performance with InfiniBand verbs-level micro-benchmarks. With this knowledge, we propose UMR-based schemes to support zero-copy datatype communication at MPI level. We show that a naive integration of UMR with an MPI stack could not bring performance benefits over existing schemes. Thus we propose two schemes -- UMR Pool and UMR Cache -- to enable high performance MPI datatype communication with UMR. To the best of our knowledge, this is the first paper to study, analyze, and design MPI noncontiguous data communication using the UMR feature. We propose and implement UMR-based designs on top of MVAPICH2 library. The experimental results at the microbenchmark level show that the proposed UMR-based design is able to deliver 4X performance improvement in latency for large message vector benchmarks over the packing/unpacking scheme. At the application level, for a 3D stencil communication kernel with MPI derived datatype on 512 processes, the optimized UMR-based design outperforms the packing/unpacking scheme by 27% in execution time.