Evaluation of three Simple Imputation Methods for Enhancing Preprocessing of Data with Missing Values

Evaluation of three Simple Imputation Methods for Enhancing Preprocessing of Data with Missing Values
复制标题

DOI:
10.5120/2619-3544
复制
发表时间:
2011-05
期刊:
International Journal of Computer Applications
影响因子:
--
通讯作者:
R. Somasundaram;R. Nedunchezhian
R. Somasundaram;R. Nedunchezhian
中科院分区:
其他
文献类型:
--
作者:
R. Somasundaram;R. Nedunchezhian

文献摘要

被引文献

相似文献

数据挖掘的一个重要阶段是预处理,即为不同的挖掘任务准备数据。通常,真实世界的数据往往是不完整的,嘈杂的,不一致的。很常见的情况是,并非每个变量的每次观测都能获得数据。因此,数据集中缺失变量的存在是显而易见的。预处理数据时最重要的任务是填充缺失值,消除噪声和纠正不一致。本文介绍了数据挖掘中的缺失值问题,并评价了一些常用的缺失值填补方法。在这项工作中,三个简单的缺失值填补方法,即(1)常数替代,(2)平均属性值替代和(3)随机属性值替代。通过使用已知的聚类方法,针对数据集中不同的缺失率或不同的缺失百分比,测量了三种缺失值填补算法的性能。为了评估性能,标准WDBC数据集已被使用。
of the important stages of data mining is preprocessing, where the data is prepared for different mining tasks. Often, the real-world data tends to be incomplete, noisy, and inconsistent. It is very common that the data are not obtainable for every observation of every variable. So the presence of missing variables is obvious in the data set. A most important task when preprocessing the data is, to fill in missing values, smooth out noise and correct inconsistencies. This paper presents the missing value problem in data mining and evaluates some of the methods generally used for missing value imputation. In this work, three simple missing value imputation methods are implemented namely (1) Constant substitution, (2) Mean attribute value substitution and (3) Random attribute value substitution. The performance of the three missing value imputation algorithms were measured with respect to different rate or different percentage of missing values in the data set by using some known clustering methods. To evaluate the performance, the standard WDBC data set has been used.