Identification of outliers in multivariate data

Identification of outliers in multivariate data
复制标题

DOI:
10.2307/2291724
复制
发表时间:
1996-09-01
影响因子:
3.7
通讯作者:
Woodruff, DL
Woodruff, DL
中科院分区:
数学1区
文献类型:
--
作者:
Rocke, DM;Woodruff, DL

文献摘要

被引文献

相似文献

对于为什么检测多元异常值的问题可能很困难以及为什么难度随着数据维度的增加而增加,给出了新的见解。描述了异常值检测方法的重大改进,并且大量的模拟实验表明混合方法扩展了异常值检测能力的实际边界。基于仿真结果和文献中的示例,研究了该算法可以检测到什么级别的污染,作为维度、计算时间、样本大小、污染分数以及污染与数据主体的距离的函数。作者和 STATLIB 提供了实现这些方法的软件。
New insights are given into why the problem of detecting multivariate outliers can be difficult and why the difficulty increases with the dimension of the data. Significant improvements in methods for detecting outliers are described, and extensive simulation experiments demonstrate that a hybrid method extends the practical boundaries of outlier detection capabilities. Based on simulation results and examples from the literature, the question of what levels of contamination can be detected by this algorithm as a function of dimension, computation time, sample size, contamination fraction, and distance of the contamination from the main body of data is investigated. Software to implement the methods is available from the authors and STATLIB.