Toward Performance Models of MPI Implementations for Understanding Application Scaling Issues

Toward Performance Models of MPI Implementations for Understanding Application Scaling Issues
复制标题

用于理解应用程序扩展问题的 MPI 实现的性能模型

DOI:
10.1007/978-3-642-15646-5_3
复制
发表时间:
2010
影响因子:
7.2
通讯作者:
J. Träff
J. Träff
中科院分区:
工程技术1区
文献类型:
--
作者:
T. Hoefler;W. Gropp;R. Thakur;J. Träff

文献摘要

被引文献

相似文献

使用MPI设计和调优并行应用程序,特别是大规模应用程序,需要了解不同算法选择和实现选项的性能含义。哪种算法更好,部分取决于不同可能的通信方法的性能,而这又取决于系统硬件和MPI实现。在缺乏针对不同MPI实现的详细性能模型的情况下,应用程序开发人员通常必须选择方法和调优代码,而无法实际估计可实现的性能并合理地为自己的选择辩护。在本文中,我们提倡构建更有用的性能模型,考虑到网络注入速率和有效二分带宽的限制。由于集体通信在实现可伸缩性方面起着至关重要的作用,因此我们还提供了集体通信算法(如broadcast、allreduce和all-to-all)的可伸缩性分析模型。我们将这些模型应用于IBM Blue Gene/P系统,并将分析性能估计值与实验测量值进行比较。
Designing and tuning parallel applications with MPI, particularly at large scale, requires understanding the performance implications of different choices of algorithms and implementation options. Which algorithm is better depends in part on the performance of the different possible communication approaches, which in turn can depend on both the system hardware and the MPI implementation. In the absence of detailed performance models for different MPI implementations, application developers often must select methods and tune codes without the means to realistically estimate the achievable performance and rationally defend their choices. In this paper, we advocate the construction of more useful performance models that take into account limitations on network-injection rates and effective bisection bandwidth. Since collective communication plays a crucial role in enabling scalability, we also provide analytical models for scalability of collective communication algorithms, such as broadcast, allreduce, and all-to-all. We apply these models to an IBM Blue Gene/P system and compare the analytical performance estimates with experimentally measured values.