MPI performance engineering with the MPI tool interface: the integration of MVAPICH and TAU

MPI performance engineering with the MPI tool interface: the integration of MVAPICH and TAU
复制标题

使用 MPI 工具接口进行 MPI 性能工程:MVAPICH 和 TAU 的集成

DOI:
--
复制
发表时间:
2017
期刊:
EuroMPI/USA
影响因子:
--
通讯作者:
D. Panda
D. Panda
中科院分区:
--
文献类型:
--
作者:
Srinivasan Ramesh;Aurèle Mahéo;S. Shende;A. Malony;H. Subramoni;D. Panda

文献摘要

被引文献

相似文献

MPI实现变得越来越复杂且高度可调,因此可伸缩性限制可能来自众多来源。作为MPI 3.0标准的一部分引入的MPI工具接口(MPI_T)为性能工具和外部软件提供了一个机会,可以在更深层次的级别进行内省和了解MPI运行时行为,以检测可伸缩性问题。该界面还提供了一种机制,可以在运行时动态重新配置MPI库以微调性能。在本文中,我们提出了一个扩展现有组件的基础架构-TAU,MVAPICH2和BEACON,以利用MPI_T接口来提供运行时内省,在线监视,建议生成和自动调整功能。我们通过开发优化生产和合成应用的优化来验证我们的设计。我们使用我们的基础架构为AmberMD [1]实施自动调整策略,该策略将MVAPICH2库内存足迹降低了20%而不会影响性能。对于集体沟通对延迟敏感的应用程序,例如Miniamr [2],我们的基础架构能够生成建议,以使MVAPICH2支持的集体卸载硬件卸载。通过实施此建议,我们看到应用程序运行时提高了5%。
MPI implementations are becoming increasingly complex and highly tunable, and thus scalability limitations can come from numerous sources. The MPI Tools Interface (MPI_T) introduced as part of the MPI 3.0 standard provides an opportunity for performance tools and external software to introspect and understand MPI runtime behavior at a deeper level to detect scalability issues. The interface also provides a mechanism to re-configure the MPI library dynamically at runtime to fine-tune performance. In this paper, we propose an infrastructure that extends existing components - TAU, MVAPICH2 and BEACON to take advantage of the MPI_T interface to offer runtime introspection, online monitoring, recommendation generation and autotuning capabilities. We validate our design by developing optimizations for a combination of production and synthetic applications. We use our infrastructure to implement an autotuning policy for AmberMD[1] that monitors and reduces MVAPICH2 library internal memory footprint by 20% without affecting performance. For applications where collective communication is latency sensitive such as MiniAMR[2], our infrastructure is able to generate recommendations to enable hardware offloading of collectives supported by MVAPICH2. By implementing this recommendation, we see a 5% improvement in application runtime.