Optimised Recovery with a Coordinated Checkpoint/Rollback Protocol for Domain Decomposition Applications

Optimised Recovery with a Coordinated Checkpoint/Rollback Protocol for Domain Decomposition Applications
复制标题

通过域分解应用程序的协调检查点/回滚协议优化恢复

DOI:
--
复制
发表时间:
2008
期刊:
Modelling, Computation and Optimization in Information Systems and Management Sciences
影响因子:
--
通讯作者:
T. Gautier
T. Gautier
中科院分区:
--
文献类型:
--
作者:
Xavier Besseron;T. Gautier

文献摘要

被引文献

相似文献

在当今长期运行的科学平行应用中,容错方案起着重要作用。由于执行过程中涉及的不可靠组件的数量,故障的可能性可能很重要。在本文中,我们介绍了基于协调方案的新检查点/回滚协议的方法和初步结果。该协议的一个功能是,由于执行的抽象表示,故障恢复仅需要部分重新启动其他过程。域分解应用程序上的模拟表明,与经典的全局回滚协议相比,重新启动所需的计算量和相关过程的数量减少。
Fault-tolerance protocols play an important role in today long runtime scientific parallel applications. The probability of a failure may be important due to the number of unreliable components involved during an execution. In this paper we present our approach and preliminary results about a new checkpoint/rollback protocol based on a coordinated scheme. One feature of this protocol is that fault recovery only requires a partial restart of other processes thanks to the availability of an abstract representation of the execution. Simulations on a domain decomposition application show that the amount of computations required to restart and the number of involved processes are reduced compared to the classical global rollback protocol.