Ensemble anomaly detection from multi-resolution trajectory features

Ensemble anomaly detection from multi-resolution trajectory features
复制标题

DOI:
10.1007/s10618-013-0334-x
复制
发表时间:
2013
影响因子:
4.8
通讯作者:
S. Ando;Theerasak Thanomphongphan;Y. Seki;Einoshin Suzuki
S. Ando;Theerasak Thanomphongphan;Y. Seki;Einoshin Suzuki
中科院分区:
计算机科学3区
文献类型:
--
作者:
S. Ando;Theerasak Thanomphongphan;Y. Seki;Einoshin Suzuki

文献摘要

被引文献

相似文献

对行为(如轨迹)的数值化、序列化观测已成为数据挖掘和知识发现研究的重要课题。将原始观察结果处理成行为的代表性特征涉及到时间尺度和分辨率的隐式选择,这严重影响了挖掘技术的最终输出。该选择与数据处理的参数相关联,例如,平滑和分割,这不直观但强烈影响数值数据的内在结构。数据挖掘技术通常需要用户提供适当处理的输入,但是选择解决方案是一项艰巨的任务,可能需要对不同设置之间的输出进行昂贵的手动检查。在本文中,我们提出了一种新的集成框架,用于在不同的尺度和分辨率参数设置下聚合异常检测任务的结果。这样的任务对于基于加权组合的现有集成方法是困难的,因为:(a)评估和加权输出需要通常不可用的异常的训练样本,(B)异常的可检测性可以取决于分辨率,即,与正常情况的区别可能仅在小的、选择性的参数范围内是明显的。在所提出的框架中,基于不同分辨率的预测聚合到行为实例的constructmeta特征表示。元特征为基于聚类的异常检测提供了判别信息。在所提出的框架中,两个相互关联的任务的行为分析:处理的数值数据和发现异常模式,共同解决,提供了一个直观的替代知识密集型的参数选择。我们还设计了一个有效的基于聚类的异常检测算法,减少了在多分辨率挖掘的计算负担。我们使用真实世界的轨迹数据进行实证研究的建议框架。它表明,所提出的框架实现了显着的改进,比传统的集成方法。
The numerical, sequential observation of behaviors, such as trajectories, have become an important subject for data mining and knowledge discovery research. Processing the raw observation into representative features of the behaviors involves an implicit choice of time-scale and resolution, which critically affect the final output of the mining techniques. The choice is associated with the parameters of data-processing, e.g., smoothing and segmentation, which unintuitively yet strongly influence the intrinsic structure of the numerical data. Data mining techniques generally require users to provide an appropriately processed input, but selecting a resolution is an arduous task that may require an expensive, manual examination of outputs between different settings. In this paper, we propose a novel ensemble framework for aggregating outcomes in different settings of scale and resolution parameters for an anomaly detection task. Such a task is difficult for existing ensemble approaches based on weighted combination because: (a) evaluating and weighing an output requires training samples of anomalies which are generally unavailable, (b) the detectability of anomalies can depend on the resolution, i.e., the distinction from normal instances may only be apparent within a small, selective range of parameters. In the proposed framework, predictions based on different resolutions are aggregated to constructmeta-feature representations of the behavior instances. Themeta-features provide the discriminative information for conducting a clustering-based anomaly detection. In the proposed framework, two interrelated tasks of the behavior analysis: processing the numerical data and discovering anomalous patterns, are addressed jointly, providing an intuitive alternative for a knowledge-intensive parameter selection. We also design an efficient clustering-based anomaly detection algorithm which reduces the computational burden of mining at multiple resolutions. We conduct an empirical study of the proposed framework using real-world trajectory data. It shows that the proposed framework achieves a significant improvement over the conventional ensemble approach.