ADAPT: an event-based adaptive collective communication framework

ADAPT: an event-based adaptive collective communication framework
复制标题

DOI:
10.1145/3208040.3208054
复制
发表时间:
2018-06
期刊:
Proceedings of the 27th International Symposium on High-Performance Parallel and Distributed Computing
影响因子:
--
通讯作者:
Xi Luo;Wei Wu;G. Bosilca;Thananon Patinyasakdikul;Linnan Wang;J. Dongarra
Xi Luo;Wei Wu;G. Bosilca;Thananon Patinyasakdikul;Linnan Wang;J. Dongarra
中科院分区:
其他
文献类型:
--
作者:
Xi Luo;Wei Wu;G. Bosilca;Thananon Patinyasakdikul;Linnan Wang;J. Dongarra

文献摘要

相似文献

高性能计算(HPC)系统的规模和异构性的增加使消息传递接口(MPI)集体通信的性能易于受到噪声的影响,并且适应硬件能力的复杂混合。最先进的MPI集合的设计严重依赖于同步;这些设计放大了参与进程之间的噪声,导致显著的性能下降。因此,这种设计理念必须重新考虑,以有效地和健壮地运行在大规模异构平台上。在本文中,我们提出了ADAPT,一个新的集体通信框架,在开放MPI,使用事件驱动技术变形集体算法异构环境。ADAPT的核心概念是放松同步,同时保持MPI集合的最小数据依赖性。为了充分利用异构系统中数据移动通道的不同带宽,我们扩展了具有拓扑感知的通信树的ADAPT集体框架。这消除了不同硬件拓扑的界限,同时最大限度地提高了数据移动的速度。我们评估我们的框架与两个流行的集体操作:广播和减少CPU和GPU集群。我们的研究结果表明,与其他最先进的MPI库相比,性能得到了大幅提高,并具有很强的抗噪声能力。特别是,我们展示了使用基于ADAPT事件的广播和减少操作的CPU数据的至少1.3倍和1.5倍加速以及GPU数据的2倍和10倍加速。
The increase in scale and heterogeneity of high-performance computing (HPC) systems predispose the performance of Message Passing Interface (MPI) collective communications to be susceptible to noise, and to adapt to a complex mix of hardware capabilities. The designs of state of the art MPI collectives heavily rely on synchronizations; these designs magnify noise across the participating processes, resulting in significant performance slowdown. Therefore, such design philosophy must be reconsidered to efficiently and robustly run on the large-scale heterogeneous platforms. In this paper, we present ADAPT, a new collective communication framework in Open MPI, using event-driven techniques to morph collective algorithms to heterogeneous environments. The core concept of ADAPT is to relax synchronizations, while mamtaining the minimal data dependencies of MPI collectives. To fully exploit the different bandwidths of data movement lanes in heterogeneous systems, we extend the ADAPT collective framework with a topology-aware communication tree. This removes the boundaries of different hardware topologies while maximizing the speed of data movements. We evaluate our framework with two popular collective operations: broadcast and reduce on both CPU and GPU clusters. Our results demonstrate drastic performance improvements and a strong resistance against noise compared to other state of the art MPI libraries. In particular, we demonstrate at least 1.3X and 1.5X speedup for CPU data and 2X and 10X speedup for GPU data using ADAPT event-based broadcast and reduce operations.