Hierarchical Clock Synchronization in MPI

Hierarchical Clock Synchronization in MPI
复制标题

MPI 中的分层时钟同步

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Conference on Cluster Computing
影响因子:
--
通讯作者:
Alexandra Carpen
Alexandra Carpen
中科院分区:
--
文献类型:
--
作者:
S. Hunold;Alexandra Carpen

文献摘要

被引文献

相似文献

MPI基准测试用于分析或调优MPI库的性能。一般来说,每个MPI库都应该根据给定的并行机进行调整,特别是在超级计算机上。系统操作员可以定义应该为特定MPI操作选择哪种算法,并且通常在分析基准测试结果之后做出选择哪种算法的决定。问题是MPI中通信操作的延迟对所选择的数据采集和数据处理方法非常敏感。出于这个原因,根据性能的测量方式,系统操作员最终可能会使用完全不同的MPI库设置。在目前的工作中,我们专注于精确测量集体行动的延迟,特别是对于小的有效载荷,外部实验因素发挥了重要作用的问题。我们提出了一种新的时钟同步算法,它利用了计算集群的层次结构,我们表明,它优于以前的方法,无论是在运行时间和精度。我们还提出了一种不同的方案,以获得精确的MPI运行时测量(称为循环时间),这是基于给定的,固定的时间片,而不是测量预定义的重复次数的传统方式。我们还强调,MPI_Barrier的使用对MPI集合的实验确定的延迟值有显着影响。我们认为,MPI_Barrier应避免,如果屏障功能的平均运行时间是在同一数量级的MPI功能的运行时间进行测量。
MPI benchmarks are used for analyzing or tuning the performance of MPI libraries. Generally, every MPI library should be adjusted to the given parallel machine, especially on supercomputers. System operators can define which algorithm should be selected for a specific MPI operation, and this decision which algorithm to select is usually made after analyzing bench-mark results. The problem is that the latency of communication operations in MPI is very sensitive to the chosen data acquisition and data processing method. For that reason, depending on how the performance is measured, system operators may end up with a completely different MPI library setup. In the present work, we focus on the problem of precisely measuring the latency of collective operations, in particular, for small payloads, where external experimental factors play a significant role. We present a novel clock synchronization algorithm, which exploits the hierarchical architecture of compute clusters, and we show that it outperforms previous approaches, both in run-time and in precision. We also propose a different scheme to obtain precise MPI run-time measurements (called Round-Time), which is based on given, fixed time slices, as opposed to the traditional way of measuring for a predefined number of repetitions. We also highlight that the use of MPI_Barrier has a significant effect on experimentally determined latency values of MPI collectives. We argue that MPI_Barrier should be avoided if the average run-time of the barrier function is in the same order of magnitude as the run-time of the MPI function to be measured.