A Flexible Checkpoint/Restart Model in Distributed Systems

A Flexible Checkpoint/Restart Model in Distributed Systems
复制标题

DOI:
10.1007/978-3-642-14390-8_22
复制
发表时间:
2009-09
期刊:
--
影响因子:
--
通讯作者:
M. Bouguerra;T. Gautier;D. Trystram;J. Vincent
M. Bouguerra;T. Gautier;D. Trystram;J. Vincent
中科院分区:
其他
文献类型:
--
作者:
M. Bouguerra;T. Gautier;D. Trystram;J. Vincent

文献摘要

被引文献

相似文献

在具有数千个处理器的新计算平台上运行的大规模应用必须面对可靠性问题。单个处理器的故障将导致整个执行失败。大多数现有的保证可靠执行的方法都基于容错机制。协调检查点是处理此类平台故障的最流行技术之一。本文提出了一种新的协调检查点/重启机制模型,适用于多种类型的计算平台。该模型的参数化过程的故障分布,成本保存一个全球一致的进程状态和计算资源的数量。通过可靠性的数学分析,我们应用这个新的模型来计算检查点时间之间的最佳间隔,以最小化平均完成时间。模型独立于失效规律的类型,使其完全灵活。我们表明,这样的模型可以用来减少检查点率高达20%,在相同的情况下,在相同的情况下,高达4倍的总开销。最后,我们报告了一些实验的基础上模拟随机故障分布对应的两个最流行的法律,即泊松过程和威布尔定律。
Large scale applications running on new computing platforms with thousands of processors have to face with reliability problems. The failure of a single processor will cause the entire execution to fail. Most existing approaches to guarantee reliable executions are based on fault tolerance mechanisms. Coordinated checkpointing is one of the most popular technique to deal with failures in such platforms. This work presents a new model of coordinated Checkpoint/Restart mechanism for several types of computing platforms. The model is parametrized by the process failure distribution, the cost to save a global consistent state of processes and the number of computational resources. Through mathematical analysis of reliability, we apply this new model to compute the optimal interval between checkpoint times in order to minimize the average completion time. Model independency from the type of the failure law makes it completely flexible. We show that such a model may be used to reduce the checkpoint rate up to 20% in same cases and up to factor 4 the total overhead in same cases. Finally, we report some experiments based on simulations for random failure distributions corresponding to the two most popular laws, namely, the Poisson’s process and Weibull’s law.