MPI datatype processing using runtime compilation

MPI datatype processing using runtime compilation
复制标题

使用运行时编译的 MPI 数据类型处理

DOI:
10.1145/2488551.2488552
复制
发表时间:
2013
期刊:
2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
T. Hoefler
T. Hoefler
中科院分区:
--
文献类型:
--
作者:
Timo Schneider;Fredrik Kjolstad;T. Hoefler

文献摘要

被引文献

相似文献

通信前后的数据打包占现代计算机通信时间的90%之多。尽管MPI为非连续数据访问提供了定义良好的数据类型接口,但出于性能原因,许多代码使用手动打包循环。程序员编写特定于访问模式的包循环(例如,手动展开),编译器为此发出优化的代码。相比之下,目前使用的MPI实现在打包时解释数据类型,导致高开销。在这项工作中,我们探讨了使用运行时编译技术在提交时为MPI数据类型生成高效和优化的包代码的有效性。因此,在打包时不会产生任何数据类型解释的开销,并且打包设置与调用函数指针一样快。我们已经实现了一个名为libpack的库,可用于编译和(un)打包MPI数据类型。该库优化了数据类型表示,并使用LLVM框架在提交时为每种数据类型生成向量化机器码。我们将展示几个示例,说明MPI数据类型包函数如何从运行时编译中获益,并分析许多应用程序中针对数据访问模式编译包函数的性能。我们表明,对于科学应用中使用的73%的数据类型,由我们的打包库生成的打包/解包函数比流行的MPI实现快7倍,并且在许多情况下优于手动打包循环。
Data packing before and after communication can make up as much as 90% of the communication time on modern computers. Despite MPI's well-defined datatype interface for non-contiguous data access, many codes use manual pack loops for performance reasons. Programmers write access-pattern specific pack loops (e.g., do manual unrolling) for which compilers emit optimized code. In contrast, MPI implementations in use today interpret datatypes at pack time, resulting in high overheads. In this work we explore the effectiveness of using runtime compilation techniques to generate efficient and optimized pack code for MPI datatypes at commit time. Thus, none of the overhead of datatype interpretation is incurred at pack time and pack setup is as fast as calling a function pointer. We have implemented a library called libpack that can be used to compile and (un)pack MPI datatypes. The library optimizes the datatype representation and uses the LLVM framework to produce vectorized machine code for each datatype at commit time. We show several examples of how MPI datatype pack functions benefit from runtime compilation and analyze the performance of compiled pack functions for the data access patterns in many applications. We show that the pack/unpack functions generated by our packing library are seven times faster than those of prevalent MPI implementations for 73% of the datatypes used in a scientific application and in many cases outperform manual pack loops.