Do We Need to Handle Every Temporal Violation in Scientific Workflow Systems?

Do We Need to Handle Every Temporal Violation in Scientific Workflow Systems?
复制标题

DOI:
10.1145/2559938
复制
发表时间:
2014-02-01
影响因子:
4.4
通讯作者:
Chen, Jinjun
Chen, Jinjun
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu, Xiao;Yang, Yun;Chen, Jinjun

文献摘要

被引文献

相似文献

科学进程通常有时间限制,有总体的最后期限和局部的里程碑。在科学工作流系统中,由于网格和云等底层计算基础设施的动态性,经常会发生执行延迟,并导致大量的时序违规。由于时间违规处理是昂贵的金钱成本和时间开销方面,一个重要的问题是“我们需要处理科学工作流系统中的每一个时间违规?根据现有的工作流时态管理的工作,答案是“真”,这些工作流时态管理采用类似于处理功能异常的哲学,即只要检测到每个时态违规,就应该处理它。然而,根据我们的观察,自我恢复的现象,执行延迟可以自动补偿的后续工作流活动的执行时间节省已被完全忽视。因此,考虑到时间违反的非功能性,我们的答案是“不一定是真的。为了利用自恢复,本文提出了一种新的自适应时间违规处理点选择策略,有效地利用了这种现象,以避免不必要的时间违规处理。基于现实世界的科学工作流和随机生成的测试用例的模拟,实验结果表明,我们的策略可以显着降低时间违规处理的成本超过96%,同时保持极低的违规率在正常情况下。
Scientific processes are usually time constrained with overall deadlines and local milestones. In scientific workflow systems, due to the dynamic nature of the underlying computing infrastructures such as grid and cloud, execution delays often take place and result in a large number of temporal violations. Since temporal violation handling is expensive in terms of both monetary costs and time overheads, an essential question aroused is "do we need to handle every temporal violation in scientific workflow systems?" The answer would be "true" according to existing works on workflow temporal management which adopt the philosophy similar to the handling of functional exceptions, that is, every temporal violation should be handled whenever it is detected. However, based on our observation, the phenomenon of self-recovery where execution delays can be automatically compensated for by the saved execution time of subsequent workflow activities has been entirely overlooked. Therefore, considering the nonfunctional nature of temporal violations, our answer is "not necessarily true." To take advantage of self-recovery, this article proposes a novel adaptive temporal violation handling point selection strategy where this phenomenon is effectively utilised to avoid unnecessary temporal violation handling. Based on simulations of both real-world scientific workflows and randomly generated test cases, the experimental results demonstrate that our strategy can significantly reduce the cost on temporal violation handling by over 96% while maintaining extreme low violation rate under normal circumstances.