AnyOut: Anytime Outlier Detection on Streaming Data

AnyOut: Anytime Outlier Detection on Streaming Data
复制标题

DOI:
10.1007/978-3-642-29038-1_18
复制
发表时间:
2012-04
期刊:
--
影响因子:
--
通讯作者:
I. Assent;P. Kranen;C. Baldauf;T. Seidl
I. Assent;P. Kranen;C. Baldauf;T. Seidl
中科院分区:
其他
文献类型:
--
作者:
I. Assent;P. Kranen;C. Baldauf;T. Seidl

文献摘要

被引文献

相似文献

随着传感器和监控应用的增加,流数据的数据挖掘越来越受到人们的关注。由于数据是不断生成的,挖掘算法需要能够以一次通过的方式分析数据。在许多应用程序中,数据对象到达的速率变化很大。这导致了用于分类或聚类的随时挖掘算法。它们成功地挖掘数据,直到流中的下一个数据的先验未知中断点。在这项工作中,我们研究了任何异常值检测。随时异常点检测是指在任何时间段内确定数据流中的对象是否异常的问题。时间越充裕,决策就越可靠。我们介绍了AnyOut,一种能够随时解决离群点检测的算法,并研究了构建底层数据结构的不同方法。我们为AnyOut提出了一个置信度度量,它可以提高恒定数据流的性能。我们在彻底的实验中评估了我们的方法,并与已有的离群值检测算法进行了比较,证明了它的性能。
With the increase of sensor and monitoring applications, data mining on streaming data is receiving increasing research attention. As data is continuously generated, mining algorithms need to be able to analyze the data in a one-pass fashion. In many applications the rate at which the data objects arrive varies greatly. This has led to anytime mining algorithms for classification or clustering. They successfully mine data until the a priori unknown point of interruption by the next data in the stream.In this work we investigate anytime outlier detection. Anytime outlier detection denotes the problem of determining within any period of time whether an object in a data stream is anomalous. The more time is available, the more reliable the decision should be. We introduce AnyOut, an algorithm capable of solving anytime outlier detection, and investigate different approaches to build up the underlying data structure. We propose a confidence measure for AnyOut that allows to improve the performance on constant data streams. We evaluate our method in thorough experiments and demonstrate its performance in comparison with established algorithms for outlier detection.