A survey of outlier detection methodologies

A survey of outlier detection methodologies
复制标题

DOI:
10.1023/b:aire.0000045502.10941.a9
复制
发表时间:
2004-10-01
影响因子:
12
通讯作者:
Austin, J
Austin, J
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hodge, VJ;Austin, J

文献摘要

被引文献

相似文献

几个世纪以来,异常值检测一直被用于检测并在适当情况下从数据中去除异常观测值。异常值的产生是由于机械故障、系统行为的变化、欺诈行为、人为错误、仪器误差,或者仅仅是由于总体中的自然偏差。对异常值的检测可以在系统故障和欺诈行为升级并可能导致灾难性后果之前将其识别出来。它可以识别错误并消除其对数据集的污染影响,从而净化数据以便进行处理。最初的异常值检测方法是随意的,但现在使用的是有原则的、系统的技术,这些技术来自计算机科学和统计学的各个领域。在本文中,我们对当代异常值检测技术进行了综述。我们确定了它们各自的动机,并在比较综述中区分了它们的优缺点。
Outlier detection has been used for centuries to detect and, where appropriate, remove anomalous observations from data. Outliers arise due to mechanical faults, changes in system behaviour, fraudulent behaviour, human error, instrument error or simply through natural deviations in populations. Their detection can identify system faults and fraud before they escalate with potentially catastrophic consequences. It can identify errors and remove their contaminating effect on the data set and as such to purify the data for processing. The original outlier detection methods were arbitrary but now, principled and systematic techniques are used, drawn from the full gamut of Computer Science and Statistics. In this paper, we introduce a survey of contemporary techniques for outlier detection. We identify their respective motivations and distinguish their advantages and disadvantages in a comparative review.