INAM: Cross-stack Profiling and Analysis of Communication in MPI-based Applications

INAM: Cross-stack Profiling and Analysis of Communication in MPI-based Applications
复制标题

INAM:基于 MPI 的应用程序中通信的跨堆栈分析和分析

DOI:
10.1145/3437359.3465582
复制
发表时间:
2021
期刊:
Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
K. Tomko
K. Tomko
中科院分区:
--
文献类型:
--
作者:
Pouya Kousha;Kamal Raj Sankarapandian Dayala Ganesh Ram;M. Kedia;H. Subramoni;Arpan Jain;A. Shafi;D. Panda;Trey Dockendorf;Heechang Na;K. Tomko

文献摘要

参考文献

被引文献

相似文献

了解HPC应用程序、MPI库、通信结构和作业调度程序之间的全栈性能权衡和相互作用是一项具有挑战性的奋进。不幸的是,现有的剖析工具是不相交的,并且仅关注剖析HPC堆栈的一个或几个级别,限制了它们可以提供的见解。在本文中,我们提出了一个标准化的方法,以促进近实时,低开销的性能表征,分析和评估的高性能通信中间件以及科学应用程序的通信使用跨堆栈的方法由INAM。根据分析会话的范围,在对应用程序进行修改和不进行修改两种模式下支持分析功能。我们使用基于MPI_T的标准化方法来设计和实现我们的设计,以获得MPI应用程序的近实时洞察力,其规模高达4,096个进程,开销小于5%。通过增加DL训练的批量大小的实验评估,我们展示了INAM实时跨栈通信分析的新优势,以检测瓶颈并解决它们,实现了用例研究的3.6倍改进。拟议的解决方案已与最新版本的INAM一起公开发布,目前正在各种HPC超级计算机上使用。
Understanding the full-stack performance trade-offs and interplay among HPC applications, MPI libraries, the communication fabric, and the job scheduler is a challenging endeavor. Unfortunately, existing profiling tools are disjoint and only focus on profiling one or a few levels of the HPC stack limiting the insights they can provide. In this paper, we propose a standardized approach to facilitate near real-time, low overhead performance characterization, profiling, and evaluation of communication of high-performance communication middleware as well as scientific applications using a cross-stack approach by INAM. The profiling capabilities are supported in two modes of with and without modifications to the application depending on the scope of the profiling session. We design and implement our designs using an MPI_T-based standardized method to obtain near real-time insights for MPI applications at scales of up to 4,096 processes with less than 5% overhead. Through experimental evaluations of increasing batch size for DL training, we demonstrate novel benefits of INAM for cross-stack communication analysis in real-time to detect bottlenecks and resolve them, achieving up to 3.6x improvements for the use-case study. The proposed solutions have been publicly released with the latest version of INAM and currently being used in production at various HPC supercomputers.
DOI: 10.1109/mcse.2015.68
发表时间: 2015
影响因子: 2.1
作者:
Palmer, Jeffrey T.;Gallo, Steven M.;Furlani, Thomas R.;Jones, Matthew D.;DeLeon, Robert L.;White, Joseph P.;Simakov, Nikolay;Patra, Abani K.;Sperhac, Jeanette;Yearke, Thomas
通讯作者: Yearke, Thomas
使用 OSU INAM 加速实时网络监控和大规模分析
DOI: 10.1145/3311790.3396672
发表时间: 2020
期刊: 2020 PEARC: Practice and Experience in Advanced Research Computing
影响因子: --
作者:
Kousha, P.;S. D., Kamal Raj;Subramoni, H.;Panda, D. K.;Na, H.;Dockendorf, T.;Tomko, K.
通讯作者: Tomko, K.
设计用于对高性能 GPU 集群进行可扩展和深入分析的分析和可视化工具
DOI: 10.1109/hipc.2019.00022
发表时间: 2019
期刊: and Analytics (HiPC
影响因子: --
作者:
Kousha, Pouya;Ramesh, Bharath;Kandadi Suresh, Kaushik;Chu, Ching-Hsiang;Jain, Arpan;Sarkauskas, Nick;Subramoni, Hari;Panda, Dhabaleswar K.
通讯作者: Panda, Dhabaleswar K.