Missing Value Imputation Based on Data Clustering

Missing Value Imputation Based on Data Clustering
复制标题

DOI:
10.1007/978-3-540-79299-4_7
复制
发表时间:
2008
期刊:
Trans. Comput. Sci.
影响因子:
--
通讯作者:
Shichao Zhang;Jilian Zhang;Xiaofeng Zhu;Yongsong Qin;Chengqi Zhang
Shichao Zhang;Jilian Zhang;Xiaofeng Zhu;Yongsong Qin;Chengqi Zhang
中科院分区:
其他
文献类型:
--
作者:
Shichao Zhang;Jilian Zhang;Xiaofeng Zhu;Yongsong Qin;Chengqi Zhang

文献摘要

被引文献

相似文献

提出了一种基于聚类的非参数缺失值插补方法CMI(Questioning-based Missing Value Imputation),用于处理目标属性中的缺失值。在我们的方法中,我们使用基于核的方法将实例A的缺失值与从不包含缺失值的实例中的数据生成的合理值进行估算,并且与实例A最相似。具体来说,我们首先将数据集(包括缺失值的实例)划分为聚类。接下来,实例A的缺失值被从A的集群中生成的合理值修补。大量的实验表明,该方法在缺失值填补任务的有效性。
We propose an efficient nonparametric missing value imputation method based on clustering, called CMI (Clustering-based Missing value Imputation), for dealing with missing values in target attributes. In our approach, we impute the missing values of an instance A with plausible values that are generated from the data in the instances which do not contain missing values and are most similar to the instance A using a kernel-based method. Specifically, we first divide the dataset (including the instances with missing values) into clusters. Next, missing values of an instance A are patched up with the plausible values generated from A’s cluster. Extensive experiments show the effectiveness of the proposed method in missing value imputation task.