Application kernels: HPC resources performance monitoring and variance analysis: HPC Resources Performance Monitoring and Variance Analysis

Application kernels: HPC resources performance monitoring and variance analysis: HPC Resources Performance Monitoring and Variance Analysis
复制标题

应用内核:HPC 资源性能监控和方差分析:HPC 资源性能监控和方差分析

DOI:
10.1002/cpe.3564
复制
发表时间:
2015
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
Patra, Abani K.
Patra, Abani K.
中科院分区:
--
文献类型:
--
作者:
Simakov, Nikolay A.;White, Joseph P.;DeLeon, Robert L.;Ghadersohi, Amin;Furlani, Thomas R.;Jones, Matthew D.;Gallo, Steven M.;Patra, Abani K.

文献摘要

相似文献

应用程序内核是计算轻量级基准或在高性能计算 (HPC) 集群上重复运行的应用程序,以便跟踪提供给用户的服务质量 (QoS)。他们成功地检测到各种硬件和软件问题,其中一些问题很严重,随后得到了纠正,从而提高了系统性能和吞吐量。在这项工作中,描述了 eXtreme Data Metrics on Demand (XDMoD) 的应用程序内核性能监控模块。通过 XDMoD 框架,应用程序内核在德克萨斯州高级计算中心的 Stampede 和 Lonestar4 集群上重复运行,总计超过 14,000 个作业。这提供了有关 HPC 集群操作的大量数据,可用于统计分析应用程序性能(通过执行时间和通信带宽等指标衡量)如何受到集群工作负载的影响。我们讨论度量分布,进行回归和相关分析,并使用 PCA 研究来描述方差并将方差与集群中应用程序的空间分布等因素相关联。最终,这些类型的分析可用于改进应用程序内核机制,从而提高交付给最终用户的 HPC 基础设施的 QoS。版权所有 © 2015 约翰·威利父子有限公司
Application kernels are computationally lightweight benchmarks or applications run repeatedly on high performance computing (HPC) clusters in order to track the Quality of Service (QoS) provided to the users. They have been successful in detecting a variety of hardware and software issues, some severe, that have subsequently been corrected, resulting in improved system performance and throughput. In this work, the application kernels performance monitoring module of eXtreme Data Metrics on Demand (XDMoD) is described. Through the XDMoD framework, the application kernels have been run repetitively on the Texas Advanced Computing Center's Stampede and Lonestar4 clusters for a total of over 14,000 jobs. This provides a body of data on the HPC clusters operation that can be used to statistically analyze how the application performance, as measured by metrics such as execution time and communication bandwidth, is affected by the cluster's workload. We discuss metric distributions, carry out regression and correlation analyses, and use a PCA study to describe the variance and relate the variance to factors such as the spatial distribution of the application in the cluster. Ultimately, these types of analyses can be used to improve the application kernel mechanism, which in turn results in improved QoS of the HPC infrastructure that is delivered to the end users. Copyright © 2015 John Wiley & Sons, Ltd.