Cache-Oblivious MPI All-to-All Communications Based on Morton Order

Cache-Oblivious MPI All-to-All Communications Based on Morton Order
复制标题

基于 Morton 顺序的缓存无关 MPI 多对多通信

DOI:
10.1109/tpds.2017.2768413
复制
发表时间:
2018-03
影响因子:
5.3
通讯作者:
Torsten Hoefler
Torsten Hoefler
中科院分区:
计算机科学2区
文献类型:
--
作者:
Shigang Li;Yunquan Zhang;Torsten Hoefler

文献摘要

参考文献

被引文献

相似文献

随着内核数量的迅速增加,众核系统对并行应用程序有效地使用其复杂的内存层次结构提出了重大挑战。许多此类应用程序在性能关键阶段依赖于集体通信,如果不进行优化,这将成为瓶颈。我们解决这个问题,提出了高速缓存无关算法MPI_Alltoall,MPI_Allgather,和MPI邻域集体利用数据的局部性。为了实现缓存无关算法,我们在共享堆上分配发送和接收缓冲区,并使用莫顿顺序来指导内存复制。我们的分析表明,我们的算法MPI_Alltoall是渐近最优的。我们展示了我们的算法的扩展,以尽量减少NUMA系统的通信距离,同时保持每个套接字内的最优性。我们进一步展示了如何缓存无关算法可以应用到多节点机。在不同的众核架构上进行了实验。对于MPI_Alltoall,我们的实现在中小型块大小的基础上实现了平均1.40倍的加速比(小于16 KB),在Xeon E7-8890上实现了平均3.03倍于MVAPICH 2的加速,在256个节点的Xeon E5-2680群集上,对于小于1 KB的数据块大小,平均实现了MVAPICH 2的2.23倍加速。
Many-core systems with a rapidly increasing number of cores pose a significant challenge to parallel applications to use their complex memory hierarchies efficiently. Many such applications rely on collective communications in performance-critical phases, which become a bottleneck if they are not optimized. We address this issue by proposing cache-oblivious algorithms for MPI_Alltoall, MPI_Allgather, and the MPI neighborhood collectives to exploit the data locality. To implement the cache-oblivious algorithms, we allocate the send and receive buffers on a shared heap and use Morton order to guide the memory copies. Our analysis shows that our algorithm for MPI_Alltoall is asymptotically optimal. We show an extension to our algorithms to minimize the communication distance on NUMA systems while maintaining optimality within each socket. We further demonstrate how the cache-oblivious algorithms can be applied to multi-node machines. Experiments are conducted on different many-core architectures. For MPI_Alltoall, our implementation achieves on average 1.40X speedup over the naive implementation based on shared heap for small and medium block sizes (less than 16 KB) on a Xeon Phi KNC, achieves on average 3.03X speedup over MVAPICH2 on a Xeon E7-8890, and achieves on average 2.23X speedup over MVAPICH2 on a 256-node Xeon E5-2680 cluster for block sizes less than 1 KB.
DOI: 10.1007/s10586-008-0065-8
发表时间: 2008-12
期刊: Cluster Computing
影响因子: --
作者:
Y. Qian;A. Afsahi
通讯作者: Y. Qian;A. Afsahi
DOI: 10.1109/ipdps.2000.846009
发表时间: 2000-02
期刊: Proceedings 14th International Parallel and Distributed Processing Symposium. IPDPS 2000
影响因子: --
作者:
N. Karonis;B. Supinski;Ian T Foster;W. Gropp;E. Lusk;J. Bresnahan
通讯作者: N. Karonis;B. Supinski;Ian T Foster;W. Gropp;E. Lusk;J. Bresnahan
DOI: 10.1002/cpe.894
发表时间: 2005-08
期刊: Concurrency and Computation: Practice and Experience
影响因子: --
作者:
Philip W. Jones;P. Worley;Y. Yoshida;James B. White;J. Levesque
通讯作者: Philip W. Jones;P. Worley;Y. Yoshida;James B. White;J. Levesque
DOI: 10.1145/377792.377895
发表时间: 2001-06
影响因子: 0.5
作者:
Hong Tang;Tao Yang
通讯作者: Hong Tang;Tao Yang
DOI: 10.1145/1810479.1810519
发表时间: 2010-06
影响因子: 0.5
作者:
G. Blelloch;Phillip B. Gibbons;H. Simhadri
通讯作者: G. Blelloch;Phillip B. Gibbons;H. Simhadri