Fast Distributed Outlier Detection in Mixed-Attribute Data Sets

Fast Distributed Outlier Detection in Mixed-Attribute Data Sets
复制标题

DOI:
10.1007/s10618-005-0014-6
复制
发表时间:
2006-05
影响因子:
4.8
通讯作者:
M. Otey;A. Ghoting;S. Parthasarathy
M. Otey;A. Ghoting;S. Parthasarathy
中科院分区:
计算机科学3区
文献类型:
--
作者:
M. Otey;A. Ghoting;S. Parthasarathy

文献摘要

被引文献

相似文献

在科学、医学和信息技术的许多领域中,有效地检测离群值或异常是一个重要的问题。应用范围从数据清理到临床诊断,从检测材料中的异常缺陷到欺诈和入侵检测。在过去的十年中,数据挖掘和统计学的研究人员已经解决了在集中设置中使用参数和非参数方法的离群值检测问题。然而,仍有若干挑战必须加以解决。首先,迄今为止,大多数方法都集中在检测连续属性空间中的离群值。然而,几乎所有真实世界的数据集都包含分类和连续属性的混合。现有的方法通常忽略或不正确地建模分类属性,导致信息的重大损失。其次,目前还没有通用的分布式离群点检测算法。大多数分布式检测算法都是针对特定领域(例如传感器网络)设计的。第三,被分析的数据集本质上可以是流式的或动态的。这样的数据集很容易出现概念漂移,数据模型也必须是动态的。为了解决这些挑战,我们提出了一个可调算法的分布式离群检测动态混合属性数据集。
Efficiently detecting outliers or anomalies is an important problem in many areas of science, medicine and information technology. Applications range from data cleaning to clinical diagnosis, from detecting anomalous defects in materials to fraud and intrusion detection. Over the past decade, researchers in data mining and statistics have addressed the problem of outlier detection using both parametric and non-parametric approaches in a centralized setting. However, there are still several challenges that must be addressed. First, most approaches to date have focused on detecting outliers in a continuous attribute space. However, almost all real-world data sets contain a mixture of categorical and continuous attributes. Categorical attributes are typically ignored or incorrectly modeled by existing approaches, resulting in a significant loss of information. Second, there have not been any general-purpose distributed outlier detection algorithms. Most distributed detection algorithms are designed with a specific domain (e.g. sensor networks) in mind. Third, the data sets being analyzed may be streaming or otherwise dynamic in nature. Such data sets are prone to concept drift, and models of the data must be dynamic as well. To address these challenges, we present a tunable algorithm for distributed outlier detection in dynamic mixed-attribute data sets.