Effect of dynamic algorithm selection of Alltoall communication on environments with unstable network speed

Effect of dynamic algorithm selection of Alltoall communication on environments with unstable network speed
复制标题

DOI:
10.1109/hpcsim.2011.5999894
复制
发表时间:
2011-07
期刊:
2011 International Conference on High Performance Computing & Simulation
影响因子:
--
通讯作者:
T. Nanri;M. Kurokawa
T. Nanri;M. Kurokawa
中科院分区:
其他
文献类型:
--
作者:
T. Nanri;M. Kurokawa

文献摘要

被引文献

相似文献

随着高性能计算系统规模的不断扩大,集体通信的性能成为一个重要的问题。通常,决定使用这些通信的哪种算法是基于消息大小和进程数量的静态指定阈值来完成的。然而,在最近使用Fat Tree或Torus拓扑作为互连的HPC系统上,网络速度变得不可预测。主要原因是争论的影响。这种效果很大程度上取决于计算节点的相对位置。另一方面,为了减少空闲节点的数量,有人尝试构建作业调度器来灵活地附加计算节点,而不考虑它们之间的相对位置。使用此策略后,网络性能会变得不稳定。为了在这样的环境下找到合适的算法,提出了一种动态方法STAR-MPI。该方法在运行时检查每个算法,并使用经验数据选择适合给定情况的算法。本文首先研究了STAR-MPI对网络速度不稳定环境的影响。在这种环境下的实验结果表明,动态方法是有效的,但测试慢算法的成本限制了效果。然后,作者提出了一种增强方法,将预测相对较慢的算法从候选列表中删除。利用算法的性能模型,结合集体通信第一次调用时测量的延迟和带宽进行预测。此时,实验结果显示的这种增强效果并不显著。然而,结果表明,通过使用更具成本效益的预测方法和优化增强中使用的阈值和因子,有可能获得更好的性能。
As the HPC systems increase their size, performance of collective communications is becoming an important issue. Usually, decisions for which algorithm of those communications to be used are done based on statically specified thresholds of the size of messages and the number of processes. However, on recent HPC systems that are hiring Fat Tree or Torus topology as their interconnect, the network speed has become unpredictable. The main reason is the effect of contentions. This effect depends heavily on the relative locations of the compute nodes. On the other hand, to reduce the number of idle nodes, there are attempts for building job schedulers to attach compute nodes flexibly, without considering their relative positions among each other. With this policy, the network performance becomes unstable. As an approach for finding an appropriate algorithm even on such environment, a dynamic method, STAR-MPI, has been proposed. This method examines each algorithm at runtime, and uses the empirical data to choose the suitable one for the given situation. This paper first examined the effect of STAR-MPI on an environment with unstable network speed. The results of experiments on this environment showed that the dynamic approach was effective, but the cost for testing slow algorithms limited the effect. Then, the authors proposed an enhancement, in which algorithms that have been predicted relatively slow were discarded from the list of candidates. The predictions were done by using the performance models of the algorithms with the latency and the bandwidth measured at the first call of the collective communication. At this point, the effect of this enhancement shown in experimental results was not significant. However, the results showed that there was a possibility for achieving better performance by using more cost-effective way of prediction and tuning thresholds and factors used in the enhancement.