System-level monitoring of floating-point performance to improve effective system utilization

System-level monitoring of floating-point performance to improve effective system utilization
复制标题

浮点性能的系统级监控以提高有效的系统利用率

DOI:
10.1145/2063348.2063355
复制
发表时间:
2011
期刊:
2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
R. Valent
R. Valent
中科院分区:
--
文献类型:
--
作者:
D. D. Vento;Thomas Engel;Siddhartha S. Ghosh;David L. Hart;Rory C. Kelly;Si Liu;R. Valent

文献摘要

被引文献

相似文献

NCAR的Bluefire超级计算机配备了一组低开销进程,可持续监控其3,840个批处理计算内核的浮点计数器。我们通过关联来自相应节点的数据来提取每个批处理作业的性能数字。从经验和良好性能的分析来看,我们使用这些数据来识别性能不佳的作业,然后与用户合作以提高其作业效率。通常,解决方案涉及简单的步骤,例如产生足够数量的进程或线程,将进程或线程绑定到核心,使用大内存页面或使用适当的编译器优化。这些努力通常会提高性能,并将挂钟运行时间减少10%到20%。随着对代码和脚本的更多更改,一些用户已经获得了40%到90%的性能改进。我们将讨论我们的仪器,一些成功的案例,以及它对其他系统的一般适用性。
NCAR's Bluefire supercomputer is instrumented with a set of low-overhead processes that continually monitor the floating point counters of its 3,840 batch-compute cores. We extract performance numbers for each batch job by correlating the data from corresponding nodes. From experience and heuristics for good performance, we use this data, in part, to identify poorly performing jobs and then work with the users to improve their job's efficiency. Often, the solution involves simple steps such as spawning an adequate number of processes or threads, binding the processes or threads to cores, using large memory pages, or using adequate compiler optimization. These efforts typically result in performance improvements and a wall-clock runtime reduction of 10% to 20%. With more involved changes to codes and scripts, some users have obtained performance improvements of 40% to 90%. We discuss our instrumentation, some successful cases, and its general applicability to other systems.