Accurate runtime selection of optimal MPI collective algorithms using analytical performance modelling

Accurate runtime selection of optimal MPI collective algorithms using analytical performance modelling
复制标题

使用分析性能建模准确选择最佳 MPI 集体算法的运行时

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
Alexey L. Lastovetsky
Alexey L. Lastovetsky
中科院分区:
--
文献类型:
--
作者:
Emin Nuriyev;Alexey L. Lastovetsky

文献摘要

被引文献

相似文献

自MPI出现以来,集体操作的性能一直是一个关键问题。许多算法已经提出了每个MPI集体操作,但没有一个证明在所有情况下都是最佳的。不同的算法表现出上级性能取决于平台,消息大小,进程的数量等MPI实现执行集体算法的选择经验,执行一个简单的运行时决策功能。虽然有效,但这种方法并不能保证最佳选择。作为一个更准确,但同样有效的替代方案,使用集体算法的选择过程的分析性能模型的建议和研究。不幸的是,以前在这方面的尝试没有成功。我们重新审视了基于分析模型的方法,并提出了两个创新,显着提高分析模型的选择精度:(1)我们从实现算法的代码,而不是从他们的高层次的数学定义来推导分析模型。这导致更详细的模型。(2)我们估计模型参数分别为每个集体的算法,并包括在相应的通信实验中执行该算法。我们的实验证明了我们的方法使用开放MPI广播和收集算法和Grid5000集群的准确性和效率。
The performance of collective operations has been a critical issue since the advent of MPI. Many algorithms have been proposed for each MPI collective operation but none of them proved optimal in all situations. Different algorithms demonstrate superior performance depending on the platform, the message size, the number of processes, etc. MPI implementations perform the selection of the collective algorithm empirically, executing a simple runtime decision function. While efficient, this approach does not guarantee the optimal selection. As a more accurate but equally efficient alternative, the use of analytical performance models of collective algorithms for the selection process was proposed and studied. Unfortunately, the previous attempts in this direction have not been successful. We revisit the analytical model-based approach and propose two innovations that significantly improve the selective accuracy of analytical models: (1) We derive analytical models from the code implementing the algorithms rather than from their high-level mathematical definitions. This results in more detailed models. (2) We estimate model parameters separately for each collective algorithm and include the execution of this algorithm in the corresponding communication experiment. We experimentally demonstrate the accuracy and efficiency of our approach using Open MPI broadcast and gather algorithms and a Grid5000 cluster.