Accelerated Real-time Network Monitoring and Profiling at Scale using OSU INAM

Accelerated Real-time Network Monitoring and Profiling at Scale using OSU INAM
复制标题

使用 OSU INAM 加速实时网络监控和大规模分析

DOI:
10.1145/3311790.3396672
复制
发表时间:
2020
期刊:
2020 PEARC: Practice and Experience in Advanced Research Computing
影响因子:
--
通讯作者:
Tomko, K.
Tomko, K.
中科院分区:
--
文献类型:
--
作者:
Kousha, P.;S. D., Kamal Raj;Subramoni, H.;Panda, D. K.;Na, H.;Dockendorf, T.;Tomko, K.

文献摘要

参考文献

被引文献

相似文献

设计一个可扩展的实时监控和分析工具,具有较低的开销,用于网络分析和自省,能够捕获所有相关的网络事件,是一项具有挑战性的任务。随着 HPC 系统变得越来越大,并且用户期望拥有更好的功能(例如细粒度的实时分析),出现了一系列新的挑战。我们通过重新设计 OSU INAM 并使其能够收集、存储、检索、可视化和分析大型复杂 HPC 集群的网络指标来应对这一挑战。增强型 OSU INAM 工具为 HPC 用户、系统管理员和 HPC 开发人员提供可扩展性、低开销和细粒度 InfiniBand 端口计数器查询和结构发现。我们的实验表明,对于包含 1,428 个节点和 114 个交换机的集群,所提出的设计可以以非常精细(亚秒)的粒度收集结构指标,并在大约 5 分钟内发现完整的网络拓扑。拟议的设计已作为 OSU INAM 工具的一部分公开发布,并可从项目网站免费下载和使用。
Designing a scalable real-time monitoring and profiling tool with low overhead for network analysis and introspection capable of capturing all relevant network events is a challenging task. Newer set of challenges come out as HPC systems are becoming larger and users are expecting to have better capabilities like real-time profiling at fine granularity. We take up this challenge by redesigning OSU INAM and making it capable to gather, store, retrieve, visualize, and analyze network metrics for large and complex HPC clusters. The enhanced OSU INAM tool provides scalability, low overhead and fined-granularity InfiniBand port counter inquiry and fabric discovery for HPC users, system administrators, and HPC developers. Our experiments show that, for a cluster of 1,428 nodes and 114 switches, the proposed design can gather fabric metrics at very fine (sub-second) granularity and discovers the complete network topology in approximately 5 minutes. The proposed design has been released publicly as a part of OSU INAM Tool and is available for free download and use from the project website.
Open MPI 中 PERUSE 接口的实现和使用
DOI: --
发表时间: 2006
期刊: PVM/MPI
影响因子: --
作者:
R. Keller;G. Bosilca;G. Fagg;Michael M. Resch;J. Dongarra
通讯作者: J. Dongarra
2019 年年度管理人:Andrea Albrecht
DOI: --
发表时间: 2019
期刊: CNE Pflegemanagement
影响因子: --
作者:
Emily R. Mackler;K. Beekman;L. Bushey;Anne Gentz;K. Davis;C. Yarrington;K. Farris;J. Griggs
通讯作者: J. Griggs
设计用于对高性能 GPU 集群进行可扩展和深入分析的分析和可视化工具
DOI: 10.1109/hipc.2019.00022
发表时间: 2019
期刊: and Analytics (HiPC
影响因子: --
作者:
Kousha, Pouya;Ramesh, Bharath;Kandadi Suresh, Kaushik;Chu, Ching-Hsiang;Jain, Arpan;Sarkauskas, Nick;Subramoni, Hari;Panda, Dhabaleswar K.
通讯作者: Panda, Dhabaleswar K.
复杂并行和分布式系统的性能技术
DOI: --
发表时间: 2000
期刊: Parallel Distributed Comput. Pract.
影响因子: --
作者:
A. Malony;S. Shende
通讯作者: S. Shende
DOI: --
发表时间: 2015
期刊: International Symposium on Computer Architecture
影响因子: --
作者:
M. Stephenson;S. Hari;Yunsup Lee;Eiman Ebrahimi;Daniel R. Johnson;D. Nellans;Mike O'Connor;S. Keckler
通讯作者: S. Keckler