Optimised Recovery with a Coordinated Checkpoint/Rollback Protocol for Domain Decomposition Applications
Optimised Recovery with a Coordinated Checkpoint/Rollback Protocol for Domain Decomposition Applications
复制标题
通过域分解应用程序的协调检查点/回滚协议优化恢复
DOI:
--
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
T. Gautier
中科院分区:
文献类型:
--
作者:
Xavier Besseron;T. Gautier
Fault-tolerance protocols play an important role in today long runtime scientific parallel applications. The probability of a failure may be important due to the number of unreliable components involved during an execution. In this paper we present our approach and preliminary results about a new checkpoint/rollback protocol based on a coordinated scheme. One feature of this protocol is that fault recovery only requires a partial restart of other processes thanks to the availability of an abstract representation of the execution. Simulations on a domain decomposition application show that the amount of computations required to restart and the number of involved processes are reduced compared to the classical global rollback protocol.