Anomaly Detection using Autoencoders in High Performance Computing Systems

Anomaly Detection using Autoencoders in High Performance Computing Systems
复制标题

DOI:
10.1609/aaai.v33i01.33019428
复制
发表时间:
2018-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Andrea Borghesi;Andrea Bartolini;M. Lombardi;M. Milano;L. Benini
Andrea Borghesi;Andrea Bartolini;M. Lombardi;M. Milano;L. Benini
中科院分区:
其他
文献类型:
--
作者:
Andrea Borghesi;Andrea Bartolini;M. Lombardi;M. Milano;L. Benini

文献摘要

被引文献

相似文献

由于超级计算机系统规模大、部件数量多,异常检测是一个非常困难的问题。目前最先进的自动异常检测技术以监督的方式使用机器学习方法或统计回归模型,这意味着检测工具经过训练以区分一组固定的行为类别(健康和不健康状态)。我们提出了一种基于机器(深度)学习技术的高性能计算系统异常检测的新方法,即一种称为自编码器的神经网络。关键思想是训练一组自动编码器来学习超级计算机节点的正常(健康)行为,并在训练后使用它们来识别异常情况。这与之前基于学习异常情况的方法不同,因为异常情况有更小的数据集(因为一开始很难识别它们)。我们在一台真正的超级计算机上测试了我们的方法,该计算机配备了细粒度、可扩展的监控基础设施,可以提供大量数据来表征系统行为。结果非常有希望:在学习正常系统行为的训练阶段之后,我们的方法能够检测到以前从未见过的异常,准确率非常高(值范围在88%到96%之间)。
Anomaly detection in supercomputers is a very difficult problem due to the big scale of the systems and the high number of components. The current state of the art for automated anomaly detection employs Machine Learning methods or statistical regression models in a supervised fashion, meaning that the detection tool is trained to distinguish among a fixed set of behaviour classes (healthy and unhealthy states).We propose a novel approach for anomaly detection in HighPerformance Computing systems based on a Machine (Deep) Learning technique, namely a type of neural network called autoencoder. The key idea is to train a set of autoencoders to learn the normal (healthy) behaviour of the supercomputer nodes and, after training, use them to identify abnormal conditions. This is different from previous approaches which where based on learning the abnormal condition, for which there are much smaller datasets (since it is very hard to identify them to begin with).We test our approach on a real supercomputer equipped with a fine-grained, scalable monitoring infrastructure that can provide large amount of data to characterize the system behaviour. The results are extremely promising: after the training phase to learn the normal system behaviour, our method is capable of detecting anomalies that have never been seen before with a very good accuracy (values ranging between 88% and 96%).