Continuous Angle-based Outlier Detection on High-dimensional Data Streams

Continuous Angle-based Outlier Detection on High-dimensional Data Streams
复制标题

DOI:
10.1145/2790755.2790775
复制
发表时间:
2015-07
期刊:
Proceedings of the 19th International Database Engineering & Applications Symposium
影响因子:
--
通讯作者:
Hao Ye;H. Kitagawa;Jun Xiao
Hao Ye;H. Kitagawa;Jun Xiao
中科院分区:
其他
文献类型:
--
作者:
Hao Ye;H. Kitagawa;Jun Xiao

文献摘要

被引文献

相似文献

数据流离群点检测是数据挖掘中一个日益重要的课题。传统的基于距离的数据流离群点检测方法不适用于高维数据集,因为在高维空间中,不同数据点之间的距离的区分度变得很差。基于角度的离群点检测是一种有效的高维离群点检测方法。本文研究了数据流上的连续ABOD问题。通常,只有少数数据对象可以在两个连续的时间戳期间改变它们的状态。因此,我们提出了几个增量的基于角度的离群值检测方法的基础上ABOD及其变种,提供了可见的速度,而不损失的准确性的数据流。首先介绍了这些增量式算法的基本思想。然后,我们解释了它们的时间复杂度。最后,我们使用合成数据流来证明他们的效率。
Outlier detection over data streams is an increasingly important task in data mining. Traditional distance-based data stream outlier detection is unsuitable for high-dimensional data sets, since the discrimination of distances between different data points becomes rather poor in high dimensional space. ABOD (Angle-based Outlier Detection) is an effective approach to detecting outliers in high-dimensional space. In this paper, the problem of continuous ABOD over data streams is studied. Generally, only a few data objects may change their states during two consecutive timestamps. Therefore, we propose several incremental angle-based outlier detection approaches over data streams based on ABOD and its variants that provide visible speed-up without loss of accuracy. Firstly, the basic ideas of these incremental algorithms are introduced. Then, we explain the time complexity of them. Finally, we use synthetic data streams to prove their efficiency.