Large scale debugging of parallel tasks with AutomaDeD

Large scale debugging of parallel tasks with AutomaDeD
复制标题

DOI:
10.1145/2063384.2063451
复制
发表时间:
2011-11
期刊:
2011 International Conference for High Performance Computing, Networking, Storage and Analysis (SC)
影响因子:
--
通讯作者:
I. Laguna;T. Gamblin;B. Supinski;S. Bagchi;G. Bronevetsky;D. Ahn;M. Schulz;B. Rountree
I. Laguna;T. Gamblin;B. Supinski;S. Bagchi;G. Bronevetsky;D. Ahn;M. Schulz;B. Rountree
中科院分区:
其他
文献类型:
--
作者:
I. Laguna;T. Gamblin;B. Supinski;S. Bagchi;G. Bronevetsky;D. Ahn;M. Schulz;B. Rountree

文献摘要

被引文献

相似文献

随着当今最大系统中内核数量的增加,开发正确的HPC应用程序仍然是一项挑战。大多数现有的调试技术在大规模上表现不佳,并且不能自动定位发生错误的并行应用程序的部分。收集大量运行时信息的开销和缺乏可扩展的错误检测算法通常会导致可扩展性差。在这项工作中,我们提出了新的,高效的技术,便于调试大规模并行应用程序的过程。我们的方法扩展了我们以前的工作,AutomaDeD,在三个主要领域以可扩展的方式隔离异常任务:(i)我们有效地比较图模型的元素(在AutomaDeD中用于对并行任务建模)使用预先计算的查找表和通过指针比较;(ii)我们在错误检测分析之前压缩每个任务的图模型,使得模型之间的比较涉及更少的元素;(iii)当出现错误和性能异常时,我们使用基于可扩展采样的聚类和最近邻技术来隔离异常任务。我们对故障注入的评估表明,AutomaDeD可以很好地扩展到数千个任务,并且可以在线方式在5秒内找到异常任务。
Developing correct HPC applications continues to be a challenge as the number of cores increases in today's largest systems. Most existing debugging techniques perform poorly at large scales and do not automatically locate the parts of the parallel application in which the error occurs. The over head of collecting large amounts of runtime information and an absence of scalable error detection algorithms generally cause poor scalability. In this work, we present novel, highly efficient techniques that facilitate the process of debugging large scale parallel applications. Our approach extends our previous work, AutomaDeD, in three major areas to isolate anomalous tasks in a scalable manner: (i) we efficiently compare elements of graph models (used in AutomaDeD to model parallel tasks) using pre-computed lookup-tables and by pointer comparison; (ii) we compress per-task graph models before the error detection analysis so that comparison between models involves many fewer elements; (iii) we use scalable sampling-based clustering and nearest-neighbor techniques to isolate abnormal tasks when bugs and performance anomalies are manifested. Our evaluation with fault injections shows that AutomaDeD scales well to thousands of tasks and that it can find anomalous tasks in under 5 seconds in an online manner.