Online Fault and Anomaly Detection for Large-Scale Scientific Workflows

Online Fault and Anomaly Detection for Large-Scale Scientific Workflows
复制标题

大规模科学工作流程的在线故障和异常检测

DOI:
--
复制
发表时间:
2011
期刊:
IEEE International Conference on High Performance Computing and Communications
影响因子:
--
通讯作者:
K. Vahi
K. Vahi
中科院分区:
--
文献类型:
--
作者:
T. Samak;D. Gunter;M. Goode;E. Deelman;G. Juve;Gaurang Mehta;Fabio Silva;K. Vahi

文献摘要

被引文献

相似文献

科学工作流程是复杂科学分析的推动者。大规模的科学工作流程在复杂的并行和分布式资源上执行,其中许多事情可能会失败。应用程序科学家需要实时跟踪工作流的状态,自动检测执行异常,并执行故障排除,而无需登录到远程节点或搜索数千个日志文件。作为nsf资助的用于归档监视性能和增强调试的综合工具(STAMPEDE)项目的一部分,我们已经开发了一个基础设施,通过集成详细的工作流和资源监视来满足这些需求。在此基础设施之上,我们开发了用于在线检测各种“硬”和“软”类型故障的分析技术。我们使用这些检测到的故障来获得关于资源和整个工作流状态的高级统计信息。在本文中,我们描述了我们的技术,并在实际应用程序日志的上下文中评估了它们的有效性。
Scientific workflows are an enabler of complex scientific analyses. Large-scale scientific workflows are executed on complex parallel and distributed resources, where many things can fail. Application scientists need to track the status of their workflows in real time, detect execution anomalies automatically, and perform troubleshooting -- without logging into remote nodes or searching through thousands of log files. As part of the NSF-funded Synthesized Tools for Archiving Monitoring Performance and Enhanced DEbugging (STAMPEDE) project, we have developed an infrastructure to answer these needs by integrating detailed workflow and resource monitoring. On top of this infrastructure, we have developed analysis techniques for online detection of a wide variety of "hard" and "soft" types of failures. We use these detected failures to derive higher-level statistics about the status of the resources and the workflow as a whole. In this paper, we describe our techniques and evaluate their effectiveness in the context of real application logs.