Sustained systems performance monitoring at the U.S. Department of Defense High Performance Computing Modernization Program

Sustained systems performance monitoring at the U.S. Department of Defense High Performance Computing Modernization Program
复制标题

美国国防部高性能计算现代化计划的持续系统性能监控

DOI:
10.1145/2063348.2063352
复制
发表时间:
2011
期刊:
2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
P. Bennett
P. Bennett
中科院分区:
--
文献类型:
--
作者:
P. Bennett

文献摘要

被引文献

相似文献

美国国防部高性能计算现代化计划(HPCMP)已经在国防部超级计算资源中心使用的高性能计算系统上实施了持续的系统性能测试。其目的是通过更新操作系统、编译器套件以及数值和通信库来监测性能改进,并监测安全补丁引起的惩罚。在实践中,每个系统的工作量是通过适当选择代表HPCMP计算技术领域的用户应用程序代码来模拟的。过去的成功案例包括Cray XT 3中OST即将出现故障、SGI Altix 4700上调度程序更新配置不完整、与Linux Networx Advanced Technology Cluster的通信库更新相关的性能问题,以及英特尔Nehalem内核从turbo模式间歇性重置为标准模式。这一历史表明,SSP测试对于向HPCMP用户提供最高质量的服务至关重要。
The U.S. Department of Defense High Performance Computing Modernization Program (HPCMP) has implemented sustained systems performance testing on high performance computing systems in use at DoD Supercomputing Resource Centers. The intent is to monitor performance improvements by updates to the operating system, compiler suites, and numerical and communications libraries, and to monitor penalties arising from security patches. In practice, each system's workload is simulated by appropriate choices of user application codes representative of the HPCMP computational technical areas. Past successes include surfacing an imminent failure of an OST in a Cray XT3, incomplete configuration of a scheduler update on an SGI Altix 4700, performance issues associated with a communications library update for a Linux Networx Advanced Technology Cluster, and intermittent resetting of Intel Nehalem cores to standard mode from turbo mode. This history demonstrates that SSP testing is critical to deliver the highest quality of service to the HPCMP users.