A Failure Prediction-Based Adaptive Checkpointing Method with Less Reliance on Temperature Monitoring for HPC Applications

A Failure Prediction-Based Adaptive Checkpointing Method with Less Reliance on Temperature Monitoring for HPC Applications
复制标题

一种基于故障预测的自适应检查点方法,较少依赖 HPC 应用的温度监控

DOI:
10.1109/cluster.2018.00067
复制
发表时间:
2018
期刊:
IEEE International Conference on Cluster Computing (CLUSTER2018)
影响因子:
--
通讯作者:
Muhammad Alfian Amrizal and Pei Li and Mulya Agung and Ryusuke Egawa and Hiroyuki Takizawa
Muhammad Alfian Amrizal and Pei Li and Mulya Agung and Ryusuke Egawa and Hiroyuki Takizawa
中科院分区:
--
文献类型:
--
作者:
Xiong Xiao;Mulya Agung;Muhammad Alfian Amrizal;Ryusuke Egawa and Hiroyuki Takizawa;Muhammad Alfian Amrizal and Pei Li and Mulya Agung and Ryusuke Egawa and Hiroyuki Takizawa

文献摘要

相似文献

恒定检查点间隔检查点方法是高性能计算领域中常用的一种检查点设置方法,已被证明是解决故障到达间隔时间服从指数分布的最优方法。另一方面,以前的作品已经表明,处理器温度和它的故障率之间有很高的相关性。通过分析并行应用程序的温度监测结果,我们注意到,故障率是动态变化的,故障到达间隔时间不遵循指数分布。在这种情况下,恒定检查点方法不是最佳解决方案,因此需要具有自适应检查点间隔的检查点方法,称为自适应检查点方法,以实现高性能。然而,使用自适应方法,处理器的温度必须不断监测,以决定检查点的时间。在本文中,我们提出了一种自适应检查点方法,对温度监测的依赖较小。我们提出的方法使用已经发生的故障(称为先前故障)的时间来估计下一次故障(称为后验故障)的平均故障时间(MTTF)。基于截断威布尔分布的特征预测后验失效的时间。仿真结果表明,该方法可以减少总浪费的时间相比,常数检查点的方法与一个相当小的温度监测周期。
Checkpointing with a constant checkpoint interval, a so-called constant checkpointing method, is commonly used in HPC field and has been proved to be the optimal solution for failures whose inter-arrival times are distributed exponentially. On the other hand, previous works have shown that there is a high correlation between processor temperature and its failure rate. By analyzing the results of the temperature monitoring on a parallel application, we noticed that the failure rate is dynamically changing and the failure inter-arrival times do not follow an exponential distribution. Under such a scenario, the constant checkpointing method is not the optimal solution and thus a checkpointing method with an adaptive checkpoint interval, called an adaptive checkpointing method, is required to achieve high performance. However, to use the adaptive method, the processor temperature must be constantly monitored in order to decide the timing for checkpointing. In this paper, we propose an adaptive checkpointing method with less reliance on the temperature monitoring. Our proposed method uses the timings of already occurred failures, called the prior failures, to estimate the mean time to failure (MTTF) of the next failure, called the posterior failure. The timing of the posterior failure is predicted based on the characteristic of a truncated Weibull distribution. The simulation results show that the proposed method can reduce the total wasted time compared to the constant checkpointing method with a considerably small temperature monitoring period.