HPC System Data Pipeline to Enable Meaningful Insights through Analysis-Driven Visualizations

HPC System Data Pipeline to Enable Meaningful Insights through Analysis-Driven Visualizations
复制标题

HPC 系统数据管道通过分析驱动的可视化提供有意义的见解

DOI:
10.1109/cluster49012.2020.00062
复制
发表时间:
2020
期刊:
2020 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
J. Brandt
J. Brandt
中科院分区:
--
文献类型:
--
作者:
B. Schwaller;Nick Tucker;Thomas W. Tucker;B. Allan;J. Brandt

文献摘要

参考文献

被引文献

相似文献

随着高性能计算(HPC)系统的复杂性不断增加,管理员和用户对系统性能和利用率的洞察力的需求也在不断增长。在HPC系统监控数据收集方面取得的进步已经产生了TB/天大小的时间序列数据集,其中包含丰富的关键信息,但是从这些指标中提取和分析有意义的信息是繁重的。我们设计并开发了一种架构,可为HPC监控数据提供灵活的、按需的运行时分析和呈现功能。我们的架构可实现快速高效的数据过滤和分析。复杂的运行时或历史分析可以表示为基于Python的计算。分析结果和各种面向HPC的摘要显示在Grafana前端界面中。为了演示我们的架构,我们将其部署到1500节点HPC系统的生产环境中,并开发了系统管理员要求的分析和可视化,随后由用户使用,以跟踪作业、用户和系统级别的集群关键指标。我们的架构是通用的,适用于任何基于 *-nix的系统,它是可扩展的,以支持多集群HPC中心。我们使用易于更换的模块来构建它,允许跨集群和中心进行独特的定制。在本文中,我们描述了数据收集和存储基础设施,创建的应用程序,以查询和分析来自自定义数据库的数据,并创建可视化显示,以提供清晰的见解HPC系统的行为。
The increasing complexity of High Performance Computing (HPC) systems has created a growing need for facilitating insight into system performance and utilization for administrators and users. The strides made in HPC system monitoring data collection have produced terabyte/day sized time-series data sets rich with critical information, but it is onerous to extract and construe meaningful information from these metrics. We have designed and developed an architecture that enables flexible, as-needed, run-time analysis and presentation capabilities for HPC monitoring data. Our architecture enables quick and efficient data filtration and analysis. Complex runtime or historical analyses can be expressed as Python-based computations. Results of analyses and a variety of HPC oriented summaries are displayed in a Grafana front-end interface. To demonstrate our architecture, we have deployed it in production for a 1500-node HPC system and have developed analyses and visualizations requested by system administrators, and later employed by users, to track key metrics about the cluster at a job, user, and system level. Our architecture is generic, applicable to any *-nix based system, and it is extensible to supporting multi-cluster HPC centers. We structure it with easily replaced modules that allow unique customization across clusters and centers. In this paper, we describe the data collection and storage infrastructure, the application created to query and analyze data from a custom database, and the visual displays created to provide clear insights into HPC system behavior.
ClusterCockpit â 用于特定作业性能监控的 Web 应用程序
DOI: 10.1109/cluster.2019.8891017
发表时间: 2019
期刊: 2019 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子: --
作者:
J. Eitzinger;T. Gruber;A. Afzal;T. Zeiser;G. Wellein
通讯作者: G. Wellein