On-line outlier detection and data cleaning

On-line outlier detection and data cleaning
复制标题

DOI:
10.1016/j.compchemeng.2004.01.009
复制
发表时间:
2004-08-15
影响因子:
4.3
通讯作者:
Jiang, W
Jiang, W
中科院分区:
工程技术2区
文献类型:
--
作者:
Liu, HC;Shah, S;Jiang, W

文献摘要

被引文献

相似文献

异常值是不遵循大量数据的统计分布的观察结果,因此可能导致统计分析的错误结果。许多传统的异常值检测工具都基于数据相同且独立分布的假设。在本文中,提出了一种抗异常数据过滤器。所提出的数据过滤器清理器包括过程模型的在线抗异常值估计,并将其与改进的卡尔曼滤波器相结合以检测和“清理”异常值。与现有方法相比,该方法具有以下特点:(a)不需要过程模型的先验知识; (b) 适用于自相关数据; (c) 可以在线实施; (d) 它尝试仅清理(即检测和替换)异常值并保留数据中的所有其他信息。 (C) 2004 Elsevier Ltd. 保留所有权利。
Outliers are observations that do not follow the statistical distribution of the bulk of the data, and consequently may lead to erroneous results with respect to statistical analysis. Many conventional outlier detection tools are based on the assumption that the data is identically and independently distributed. In this paper, an outlier-resistant data filter-cleaner is proposed. The proposed data filter-cleaner includes an on-line outlier-resistant estimate of the process model and combines it with a modified Kalman filter to detect and "clean" outliers. The advantage over existing methods is that the proposed method has the following features: (a) a priori knowledge of the process model is not required; (b) it is applicable to autocorrelated data; (c) it can be implemented on-line; and (d) it tries to only clean (i.e., detects and replaces) outliers and preserves all other information in the data. (C) 2004 Elsevier Ltd. All rights reserved.