Online Self-Evolving Anomaly Detection for Reliable Cloud Computing

Online Self-Evolving Anomaly Detection for Reliable Cloud Computing
复制标题

DOI:
10.1109/ucc56403.2022.00014
复制
发表时间:
2022-12
期刊:
2022 IEEE/ACM 15th International Conference on Utility and Cloud Computing (UCC)
影响因子:
--
通讯作者:
Tianyu Bai;Haili Wang;Jingda Guo;Xu Ma;Mahendra Talasila;Sihai Tang;Song Fu;Qing Yang
Tianyu Bai;Haili Wang;Jingda Guo;Xu Ma;Mahendra Talasila;Sihai Tang;Song Fu;Qing Yang
中科院分区:
其他
文献类型:
--
作者:
Tianyu Bai;Haili Wang;Jingda Guo;Xu Ma;Mahendra Talasila;Sihai Tang;Song Fu;Qing Yang

文献摘要

相似文献

生产云计算系统由数百到数千个计算和存储节点组成。如此规模,加上不断增长的系统复杂性,对可靠云计算的故障和资源管理造成了重大挑战。高效的系统监控和故障检测对于理解突发的云范围现象和智能管理云资源以实现系统级可靠性保证和应用程序级性能保证至关重要。为了检测故障,我们需要监控云执行并收集运行时性能数据。在现实系统中,这些数据在运行时通常是未标记的,因此先前的故障历史记录并不总是可用。在本文中,我们提出了一种用于云可靠性保证的自我进化异常检测框架。我们的框架不需要任何先前的故障历史记录,并且它通过不断探索新验证的异常记录并在运行时不断更新异常检测器来自我进化,而无需昂贵的模型重新训练。我们框架的一个明显优势是,云系统操作员只需要检查少量检测到的异常(与数以亿计的系统/应用程序事件记录相比),并且他们的决策可用于更新检测器。因此,检测器随着系统硬件的升级、软件堆栈的更新和用户工作负载的变化而发展。此外,我们设计了两种类型的检测器,一种用于一般异常检测,另一种用于特定类型的异常检测。利用自我进化和在线学习技术,我们的检测器平均可以达到 88.94% 的灵敏度和 94.60% 的特异性,这使得它们适合实际部署。
Production cloud computing systems consist of hundreds to thousands of computing and storage nodes. Such a scale, combined with ever-growing system complexity, is causing a key challenge to failure and resource management for dependable cloud computing. Efficient system monitoring and failure detection are crucial for understanding emergent, cloudwide phenomena and intelligently managing cloud resources for system-level dependability assurance and application-level performance assurance. To detect failures, we need to monitor the cloud execution and collect runtime performance data. These data are usually unlabeled at runtime in real-world systems, and thus a prior failure history is not always available. In this paper, we present a self-evolving anomaly detection framework for cloud dependability assurance. Our framework does not require any prior failure history, and it self-evolves by continuously exploring newly verified anomaly records and continuously updating the anomaly detector at runtime without expensive model retraining. A distinct advantage of our framework is that cloud system operators only need to check a small number of detected anomalies (compared with thousands-millions of system/application event records) and their decisions are leveraged to update the detector. Thus, the detector evolves following the upgrade of system hardware, update of software stack, and change of user workloads. Moreover, we design two types of detectors, one for general anomaly detection and the other for type-specific anomaly detection. Leveraging self-evolution and online learning techniques, our detectors can achieve 88.94% of sensitivity and 94.60% of specificity on average, which makes them suitable for real-world deployment.