Failure analysis of distributed scientific workflows executing in the cloud

Failure analysis of distributed scientific workflows executing in the cloud
复制标题

在云中执行的分布式科学工作流程的故障分析

DOI:
--
复制
发表时间:
2012
期刊:
2012 8th international conference on network and service management (cnsm) and 2012 workshop on systems virtualiztion management (svm)
影响因子:
--
通讯作者:
K. Vahi
K. Vahi
中科院分区:
--
文献类型:
--
作者:
T. Samak;D. Gunter;M. Goode;E. Deelman;G. Juve;Fabio Silva;K. Vahi

文献摘要

被引文献

相似文献

这项工作提出了描述在Amazon EC2上执行大型科学应用程序期间观察到的故障的模型。科学工作流被用作应用程序表示的底层抽象。随着科学工作流程扩展到成千上万个不同的任务,由于软件和硬件故障导致的故障变得越来越普遍。我们通过Stampede框架研究了从4个科学应用中收集的数据的工作失败模型。特别地,我们证明了朴素贝叶斯分类器可以准确地预测作业的失效概率。这些模型允许我们预测给定执行资源的作业失败,然后将这些失败预测用于两个更高级别的目标:(1)建议更好的作业分配,(2)向工作流组件开发人员提供关于其应用程序代码健壮性的定量反馈。
This work presents models characterizing failures observed during the execution of large scientific applications on Amazon EC2. Scientific workflows are used as the underlying abstraction for application representations. As scientific workflows scale to hundreds of thousands of distinct tasks, failures due to software and hardware faults become increasingly common. We study job failure models for data collected from 4 scientific applications, by our Stampede framework. In particular, we show that a Naive Bayes classifier can accurately predict the failure probability of jobs. The models allow us to predict job failures for a given execution resource and then use these failure predictions for two higher-level goals: (1) to suggest a better job assignment, and (2) to provide quantitative feedback to the workflow component developer about the robustness of their application codes.