STAR-MPI: self tuned adaptive routines for MPI collective operations

STAR-MPI: self tuned adaptive routines for MPI collective operations
复制标题

DOI:
10.1145/1183401.1183431
复制
发表时间:
2006-06
期刊:
--
影响因子:
--
通讯作者:
Ahmad Faraj;Xin Yuan;D. Lowenthal
Ahmad Faraj;Xin Yuan;D. Lowenthal
中科院分区:
其他
文献类型:
--
作者:
Ahmad Faraj;Xin Yuan;D. Lowenthal

文献摘要

被引文献

相似文献

消息传递接口(MPI)集体通信例程被广泛用于并行应用程序中。为了制定集体通信程序以在不同平台上的不同应用程序实现高性能,必须适应系统体系结构和应用程序工作负载。当前的MPI实现不支持这种软件适应性,并且无法在许多平台上实现高性能。在本文中,我们介绍了Star MPI(MPI集体操作的自调自适应例程),这是一组MPI集体通信例程,能够适应系统体系结构和应用程序工作负载。对于每个操作,Star-MPI都保持了一组通信算法,这些算法在不同情况下可能有效。当应用程序执行时,STAR-MPI例程在运行时应用了软件(AEOS)技术的自动经验优化,以动态选择平台上应用程序的最佳性能算法。我们描述了Star-MPI中使用的技术,分析Star-MPI开销,并通过应用和基准评估Star-MPI的性能。我们的研究结果表明,Star-MPI是强大而有效的。它能够具有合理的开销和有效的算法,并且在许多情况下,它在很大程度上超过了传统的MPI实现。
Message Passing Interface (MPI) collective communication routines are widely used in parallel applications. In order for a collective communication routine to achieve high performance for different applications on different platforms, it must be adaptable to both the system architecture and the application workload. Current MPI implementations do not support such software adaptability and are not able to achieve high performance on many platforms. In this paper, we present STAR-MPI (Self Tuned Adaptive Routines for MPI collective operations), a set of MPI collective communication routines that are capable of adapting to system architecture and application workload. For each operation, STAR-MPI maintains a set of communication algorithms that can potentially be efficient at different situations. As an application executes, a STAR-MPI routine applies the Automatic Empirical Optimization of Software (AEOS) technique at run time to dynamically select the best performing algorithm for the application on the platform. We describe the techniques used in STAR-MPI, analyze STAR-MPI overheads, and evaluate the performance of STAR-MPI with applications and benchmarks. The results of our study indicate that STAR-MPI is robust and efficient. It is able to and efficient algorithms with reasonable overheads, and it out-performs traditional MPI implementations to a large degree in many cases.