Imputation of missing data in industrial databases

Imputation of missing data in industrial databases
复制标题

DOI:
10.1023/a:1008334909089
复制
发表时间:
1999-11-01
影响因子:
5.3
通讯作者:
Samad, T
Samad, T
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lakshminarayan, K;Harp, SA;Samad, T

文献摘要

被引文献

相似文献

IDA方法在许多领域应用的一个限制因素是数据存储库的不完整性。许多记录都有未填写的字段,特别是在手动输入数据时。此外,条目的很大一部分可能是错误的,并且除了丢弃这些记录之外可能没有其他选择。但数据库中的每个单元并不是一个独立的数据。统计关系将约束并通常决定缺失值。因此,数据估算,即填补部分缺失数据的缺失值,可以成为许多国际开发协会项目的宝贵的第一步。新的填补方法,可以处理大规模的问题和大规模稀疏的工业数据库是必要的。为了说明不完整的数据库问题,我们分析了一个数据库与仪表维护和测试记录的工业过程。尽管有工艺数据收集的法规要求,但该数据库的完整性不到50%。接下来,我们将讨论缺失数据问题的可能解决方案。几种方法来填补指出,并分为两类:数据驱动和基于模型。然后,我们描述了我们使用过的两种基于机器学习的方法。这些都建立在众所周知的算法之上:AutoClass和C4.5。设计了几个实验,所有使用的维护数据库作为一个共同的测试床,但与各种数据分割和算法的变化。结果通常为阳性,插补准确率高达80%。最后,我们概述了在选择插补方法的一些考虑,并通过讨论智能数据分析的数据插补的应用文件。
A limiting factor for the application of IDA methods in many domains is the incompleteness of data repositories. Many records have fields that are not filled in, especially, when data entry is manual. In addition, a significant fraction of the entries can be erroneous and there may be no alternative but to discard these records. But every cell in a database is not an independent datum. Statistical relationships will constrain and, often determine, missing values. Data imputation, the filling in of missing values for partially missing data, can thus be an invaluable first step in many IDA projects. New imputation methods that can handle the large-scale problems and large-scale sparsity of industrial databases are needed. To illustrate the incomplete database problem, we analyze one database with instrumentation maintenance and test records for an industrial process. Despite regulatory requirements for process data collection, this database is less than 50% complete. Next, we discuss possible solutions to the missing data problem. Several approaches to imputation are noted and classified into two categories: data-driven and model-based. We then describe two machine-learning-based approaches that we have worked with. These build upon well-known algorithms: AutoClass and C4.5. Several experiments are designed, all using the maintenance database as a common test-bed but with various data splits and algorithmic variations. Results are generally positive with up to 80% accuracies of imputation. We conclude the paper by outlining some considerations in selecting imputation methods, and by discussing applications of data imputation for intelligent data analysis.