NWPerf: a system wide performance monitoring tool for large Linux clusters

NWPerf: a system wide performance monitoring tool for large Linux clusters
复制标题

NWPerf:大型 Linux 集群的系统范围性能监控工具

DOI:
10.1109/clustr.2004.1392637
复制
发表时间:
2004
期刊:
2004 IEEE International Conference on Cluster Computing (IEEE Cat. No.04EX935)
影响因子:
--
通讯作者:
J. Nieplocha
J. Nieplocha
中科院分区:
--
文献类型:
--
作者:
Ryan W. Mooney;R. S. Studham;Kenneth P Schmidt;J. Nieplocha

文献摘要

被引文献

相似文献

我们推出了 NWPerf,这是一种用于分析大规模超级计算集群上细粒度性能指标数据的新系统。该工具能够从全局系统角度测量整个系统的应用程序效率,并提供各个应用程序的详细视图。 NWPerf 提供此服务,同时最大限度地减少对用户应用程序性能的影响。我们描述了可以从系统中获取的信息类型,并演示了如何使用该系统来检测和消除应用程序中的性能问题,从而将性能提高高达数千%。 NWPerf 架构已被证明是一个稳定且可扩展的平台,用于在 PNNL 的大型 1954-CPU 生产 Linux 集群上收集性能数据。
We present NWPerf, a new system for analyzing fine granularity performance metric data on large-scale supercomputing clusters. This tool is able to measure application efficiency on a system wide basis from both a global system perspective as well as providing a detailed view of individual applications. NWPerf provides this service while minimizing the impact on the performance of user applications. We describe the type of information that can be derived from the system, and demonstrate how the system was used detect and eliminate a performance problem in an application application that improved performance by up to several thousand percent. The NWPerf architecture has proven to be a stable and scalable platform for gathering performance data on a large 1954-CPU production Linux cluster at PNNL.