Aarohi: Making Real-Time Node Failure Prediction Feasible

Aarohi: Making Real-Time Node Failure Prediction Feasible
复制标题

DOI:
10.1109/ipdps47924.2020.00115
复制
发表时间:
2020-05
期刊:
2020 IEEE International Parallel and Distributed Processing Symposium (IPDPS)
影响因子:
--
通讯作者:
Anwesha Das;F. Mueller;B. Rountree
Anwesha Das;F. Mueller;B. Rountree
中科院分区:
其他
文献类型:
--
作者:
Anwesha Das;F. Mueller;B. Rountree

文献摘要

被引文献

相似文献

众所周知,大规模生产系统会遇到节点故障,这会影响计算能力和能源。在HPC系统和企业数据中心中,随着硬件和软件复杂性的增加,应对故障变得越来越具有挑战性。在这种系统中的异常检测的背景下,日志的几个数据挖掘解决方案进行了研究。然而,随着后续的主动故障缓解,现有的日志挖掘解决方案对于实时异常检测来说不够快。基于机器学习(ML)的训练可以产生高精度,但推理方案需要通过快速解析器来增强,以实时评估异常。本文提出了一个基于上下文无关文法的快速事件分析框架Aarohi1,它描述了一种有效的在线故障预测方法。Aarohi被设计为通用和可扩展的,使其适合作为实时预测器。Aarohi获得了超过3分钟的节点故障提前时间,平均预测时间为0.31毫秒,链长为18。获得的总体改善w.r.t.现有技术水平超过27.4倍。我们的基于编译器的方法提供了新的研究方向,提前期优化与一个显着的预测加速所需的主动容错解决方案的部署在实践中。
Large-scale production systems are well known to encounter node failures, which affect compute capacity and energy. Both in HPC systems and enterprise data centers, combating failures is becoming challenging with increasing hardware and software complexity. Several data mining solutions of logs have been investigated in the context of anomaly detection in such systems. However, with subsequent proactive failure mitigation, the existing log mining solutions are not sufficiently fast for real-time anomaly detection. Machine learning (ML)-based training can produce high accuracy but the inference scheme needs to be enhanced with rapid parsers to assess anomalies in real-time. This work tackles online anomaly prediction in computing systems by exploiting context free grammar-based rapid event analysis.We present our framework Aarohi1, which describes an effective way to predict failures online. Aarohi is designed to be generic and scalable making it suitable as a real-time predictor. Aarohi obtains more than 3 minutes lead times to node failures with an average of 0.31 msecs prediction time for a chain length of 18. The overall improvement obtained w.r.t. the existing state-of-the-art is over a factor of 27.4×. Our compiler-based approach provides new research directions for lead time optimization with a significant prediction speedup required for the deployment of proactive fault tolerant solutions in practice.