Performance analysis of MPI collective operations

Performance analysis of MPI collective operations
复制标题

DOI:
10.1007/s10586-007-0012-0
复制
发表时间:
2007-06-01
影响因子:
4.4
通讯作者:
Dongarra, Jack J.
Dongarra, Jack J.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Pjesivac-Grbovic, Jelena;Angskun, Thara;Dongarra, Jack J.

文献摘要

被引文献

相似文献

以前对应用程序使用的研究表明,集体通信的性能对高性能计算至关重要。尽管该领域的研究很活跃。对于集体沟通问题的优化,目前还缺乏通用的、可行的解决方案。在本文中,我们通过将公认的点对点通信模型(如Hockney、LogP/LogGP和PLogP)扩展到集体操作,分析并尝试在广泛部署的MPI编程范式的背景下改进集群内集体通信。我们将模型预测与实验数据进行比较,并利用这些结果构建广播集体的最优决策函数。我们定量地比较了基于模型的决策函数与实验最优决策函数的质量。此外,在这项工作中,我们还介绍了一种新的优化的基于树的广播算法,分割二进制。我们的结果表明,所有的模型都可以为不同算法的各个方面以及它们的相对性能提供有用的见解。尽管如此,根据我们的发现,我们认为完全依赖模型不会产生最佳结果。此外,我们的实验结果已经确定了间隙参数对于经典的点对点管道和我们对扇形拓扑的扩展的精确建模是最关键的。
Previous studies of application usage show that the performance of collective communications are critical for high-performance computing. Despite active research in the field. both general and feasible solution to the optimization of collective communication problem is still missing.In this paper, we analyze and attempt to improve intracluster collective communication in the context of the widely deployed MPI programming paradigm by extending accepted models of point-to-point communication, such as Hockney, LogP/LogGP, and PLogP, to collective operations. We compare the predictions from models against the experimentally gathered data and using these results, construct optimal decision function for broadcast collective. We quantitatively compare the quality of the model-based decision functions to the experimentally-optimal one. Additionally, in this work, we also introduce a new form of an optimized tree-based broadcast algorithm, splitted-binary.Our results show that all of the models can provide useful insights into various aspects of the different algorithms as well as their relative performance. Still, based on our findings, we believe that the complete reliance on models would not yield optimal results. In addition, our experimental results have identified the gap parameter as being the most critical for accurate modeling of both the classical point-to-point-based pipeline and our extensions to fan-out topologies.