A Year of Automated Anomaly Detection in a Datacenter

A Year of Automated Anomaly Detection in a Datacenter
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
--
影响因子:
--
通讯作者:
Rufaida Ahmed;J. Porter;Abubaker Abdelmutalab;R. Ricci
Rufaida Ahmed;J. Porter;Abubaker Abdelmutalab;R. Ricci
中科院分区:
其他
文献类型:
--
作者:
Rufaida Ahmed;J. Porter;Abubaker Abdelmutalab;R. Ricci

文献摘要

相似文献

基于机器学习的异常检测可以成为理解大型复杂计算机系统行为的强大工具。然而,所看到的异常集可能会随着时间的推移而变化:随着系统的发展,被投入不同的用途,并遇到不同的工作负载,其“典型”行为和它遇到的异常也可能会发生变化。这自然提出了两个问题:在这种情况下,自动异常检测的有效性如何,以及异常行为随着时间的推移会发生多大变化?在本文中,我们研究了这些问题的数据集,从一个系统,管理生命周期的服务器在数据中心。我们查看了一个大约有500台服务器的数据中心一年的运行日志。应用最先进的技术来发现异常事件,我们发现有一组“核心”异常模式在整个研究期间持续存在,但是为了跟踪系统的演变,我们必须定期重新训练检测器。与该系统的管理员合作,我们发现,尽管模式发生了这些变化,但它们仍然包含可操作的见解。
—Anomaly detection based on Machine Learning can be a powerful tool for understanding the behavior of large, complex computer systems in the wild. The set of anomalies seen, however, can change over time: as the system evolves, is put to different uses, and encounters different workloads, both its ‘typical’ behavior and the anomalies that it encounters can change as well. This naturally raises two questions: how effective is automated anomaly detection in this setting, and how much does anomalous behavior change over time? In this paper, we examine these question for a dataset taken from a system that manages the lifecycle of servers in datacenters. We look at logs from one year of operation of a datacenter of about 500 servers. Applying state-of-the art techniques for finding anomalous events, we find that there are a ‘core’ set of anomaly patterns that persist over the entire period studied, but that in to track the evolution of the system, we must re-train the detector periodically. Working with the administrators of this system, we find that, despite these changes in patterns, they still contain actionable insights.