International Journal of High Performance Computing Applications Failure Prediction for Hpc Systems and Applications: Current Situation and Open Issues Failure Prediction for Hpc Systems and Applications: Current Situation and Open Issues

International Journal of High Performance Computing Applications Failure Prediction for Hpc Systems and Applications: Current Situation and Open Issues Failure Prediction for Hpc Systems and Applications: Current Situation and Open Issues
复制标题

国际高性能计算应用杂志 HPC 系统和应用的故障预测:现状和未解决的问题 HPC 系统和应用的故障预测:当前的情况和未解决的问题

DOI:
--
复制
发表时间:
--
期刊:
--
影响因子:
--
通讯作者:
W. Kramer
W. Kramer
中科院分区:
--
文献类型:
--
作者:
Ana Gainaru;F. Cappello;M. Snir;W. Kramer

文献摘要

被引文献

相似文献

随着大规模系统向后千万亿次计算发展,重点关注提供旨在最小化故障对应用程序影响的容错策略至关重要。到目前为止,最流行的技术是检查点重启策略。这种经典方法的一个补充是故障避免,通过它可以预测故障的发生并采取主动措施。这需要一个可靠的预测系统来预测故障及其位置。提供预测的一种方法是分析大型系统在生产过程中生成的系统日志。目前在这一领域的研究提出了一些限制,使他们无法运行在真实的生产高性能计算(HPC)系统。基于我们的观察,不同的故障有不同的分布和行为,我们提出了一种新的混合方法,结合信号分析与数据挖掘,以克服目前的局限性。我们表明,通过分析每个事件,根据其特定的行为,我们的预测提供了超过90%的精度,它能够发现约50%的系统中的所有故障,结果,允许其集成在主动容错协议。
As large-scale systems evolve towards post-petascale computing, it is crucial to focus on providing fault-tolerance strategies that aim to minimize fault's effects on applications. By far the most popular technique is the checkpoint–restart strategy. A complement to this classical approach is failure avoidance, by which the occurrence of a fault is predicted and proactive measures are taken. This requires a reliable prediction system to anticipate failures and their locations. One way of offering prediction is by the analysis of system logs generated during production by large-scale systems. Current research in this field presents a number of limitations that make them unusable for running on real production high-performance computing (HPC) systems. Based on our observations that different failures have different distributions and behaviours, we propose a novel hybrid approach that combines signal analysis with data mining in order to overcome current limitations. We show that by analysing each event according to its specific behaviour, our prediction provides a precision of over 90% and its able to discover about 50% of all failures in a system, result which allows its integration in proactive fault tolerance protocols.