Comprehensive job level resource usage measurement and analysis for XSEDE HPC systems

Comprehensive job level resource usage measurement and analysis for XSEDE HPC systems
复制标题

XSEDE HPC 系统的全面作业级别资源使用情况测量和分析

DOI:
10.1145/2484762.2484781
复制
发表时间:
2013
期刊:
Proceedings of the Conference on Extreme Science and Engineering Discovery Environment: Gateway to Discovery (XSEDE '13
影响因子:
--
通讯作者:
Patra, Abani K.
Patra, Abani K.
中科院分区:
--
文献类型:
--
作者:
Lu, Charng-Da;Browne, James;DeLeon, Robert L.;Hammond, John;Barth, William;Furlani, Thomas R.;Gallo, Steven M.;Jones, Matthew D.;Patra, Abani K.

文献摘要

参考文献

被引文献

相似文献

本文提出了一种方法,全面的作业级资源使用的测量和分析和应用程序的分析,规划HPC系统和案例研究应用程序的方法,XSEDE游侠和Lonestar4系统在得克萨斯大学。该方法的步骤是:在全系统收集作业和节点一级的资源使用和业绩统计数据,将所产生的作业数据映射和储存到关系数据库,这便于进一步执行和将数据转换为具体统计和分析算法所需的格式。分析可以在不同的粒度级别执行:作业、用户或系统范围的基础。测量基于一个新的轻量级的以作业为中心的测量工具“TACC_Stats”[1],它收集了所有计算节点上的一组全面的指标。数据映射和分析工具将是XSEDE社区的XDMoD项目[2]的扩展。本文还报告了德克萨斯州高级计算中心的Lonestar4和Ranger超级计算机的测量数据分析的初步结果。所介绍的案例研究表明,当TACC_Stats在整个XSEDE系统中部署时,所有资源都可以获得详细信息。该方法可以应用于任何运行TACC_Stats度量工具的系统。
This paper presents a methodology for comprehensive job level resource use measurement and analysis and applications of the analyses to planning for HPC systems and a case study application of the methodology to the XSEDE Ranger and Lonestar4 systems at the University of Texas. The steps in the methodology are: System-wide collection of resource use and performance statistics at the job and node levels, mapping and storage of the resultant job-wise data to a relational database which eases further implementation and transformation of data to the formats required by specific statistical and analytical algorithms. Analyses can be carried out at different levels of granularity: job, user, or system-wide basis. Measurements are based on a novel lightweight job-centric measurement tool "TACC_Stats" [1], which gathers a comprehensive set of metrics on all compute nodes. The data mapping and analysis tools will be an extension to the XDMoD project [2] for the XSEDE community. This paper also reports the preliminary results from the analysis of measured data for Texas Advanced Computing Center's Lonestar4 and Ranger supercomputers. The case studies presented indicate the level of detailed information that will be available for all resources when TACC_Stats is deployed throughout the XSEDE system. The methodology can be applied to any system that runs the TACC_Stats measurement tool.
浮点性能的系统级监控以提高有效的系统利用率
DOI: 10.1145/2063348.2063355
发表时间: 2011
期刊: 2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子: --
作者:
D. D. Vento;Thomas Engel;Siddhartha S. Ghosh;David L. Hart;Rory C. Kelly;Si Liu;R. Valent
通讯作者: R. Valent
进一步提高 Scalasca 工具集的可扩展性
DOI: 10.1007/978-3-642-28145-7_45
发表时间: 2010
期刊: 2010 ACM/IEEE International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子: --
作者:
M. Geimer;P. Saviankou;A. Strube;Z. Szebenyi;F. Wolf;B. Wylie
通讯作者: B. Wylie
HPCToolkit:科学计算的性能工具
DOI: 10.1088/1742-6596/125/1/012088
发表时间: 2008
期刊: Journal of Physics: Conference Series
影响因子: --
作者:
Nathan R. Tallent;J. Mellor;L. Adhianto;M. Fagan;Mark W. Krentel
通讯作者: Mark W. Krentel
使用高性能计算机系统的应用程序内核的性能指标和审核框架
DOI: 10.1002/cpe.2871
发表时间: 2013
期刊: Concurrency and Computation: Practice and Experience
影响因子: --
作者:
Furlani, Thomas R.;Jones, Matthew D.;Gallo, Steven M.;Bruno, Andrew E.;Lu, Charng‐Da;Ghadersohi, Amin;Gentner, Ryan J.;Patra, Abani;DeLeon, Robert L.;Laszewski, Gregor
通讯作者: Laszewski, Gregor