Missing Value Imputation Based on Data Clustering
Missing Value Imputation Based on Data Clustering
复制标题
DOI:
10.1007/978-3-540-79299-4_7
复制
发表时间:
2008
期刊:
影响因子:
--
通讯作者:
Shichao Zhang;Jilian Zhang;Xiaofeng Zhu;Yongsong Qin;Chengqi Zhang
中科院分区:
文献类型:
--
作者:
Shichao Zhang;Jilian Zhang;Xiaofeng Zhu;Yongsong Qin;Chengqi Zhang
We propose an efficient nonparametric missing value imputation method based on clustering, called CMI (Clustering-based Missing value Imputation), for dealing with missing values in target attributes. In our approach, we impute the missing values of an instance A with plausible values that are generated from the data in the instances which do not contain missing values and are most similar to the instance A using a kernel-based method. Specifically, we first divide the dataset (including the instances with missing values) into clusters. Next, missing values of an instance A are patched up with the plausible values generated from A’s cluster. Extensive experiments show the effectiveness of the proposed method in missing value imputation task.