Efficient Utilization of Missing Data in Cost-Sensitive Learning

Efficient Utilization of Missing Data in Cost-Sensitive Learning
复制标题

在成本敏感型学习中有效利用缺失数据

DOI:
10.1109/tkde.2019.2956530
复制
发表时间:
2019-11
影响因子:
8.9
通讯作者:
Zhang Shichao
Zhang Shichao
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zhu Xiaofeng;Yang Jianye;Zhang Chengyuan;Zhang Shichao

文献摘要

参考文献

被引文献

相似文献

与以往填补方法利用完全样本中的信息填补不完全样本中的缺失值不同,本文提出了一种日期驱动增量填补模型(DIM),它利用数据集中的所有可用信息,经济、有效、有序、迭代地填补缺失值。为此,我们提出了一个评分规则,考虑到经济标准和有效的填补信息的缺失功能进行排名。经济性准则同时考虑了特征的插补成本和判别能力,而有效插补信息使得能够利用数据集中所有观测信息(包括插补的缺失值)来插补剩余的缺失值。在插补过程中,我们的DIM首先检测不需要插补的样本,以减少插补成本和噪声,然后选择排名靠前的缺失特征首先进行插补。插补过程依次插补缺失特征,直到所有缺失值都被插补或插补成本耗尽。UCI数据集上的实验结果表明,我们提出的DIM的优势,比较方法,在预测精度和分类精度。
Different from previous imputation methods which impute missing values in the incomplete samples by using the information in the complete samples, this paper proposes a Date-drive Incremental imputation Model, DIM for short, which uses all available information in the data set to impute missing values economically, effectively, orderly, and iteratively. To this end, we propose a scoring rule to rank the missing features by taking into account both the economical criterion and the effective imputation information. The economical criterion takes both the imputation cost and the discriminative ability of the feature into account, while the effective imputation information enables to use all observed information in the data set including the imputed missing values to impute the left missing values. During the imputation process, our DIM first detects the neednot-impute samples for reducing the imputation cost and noise, and then selects the missing features with the top rank to impute first. The imputation process orderly imputes the missing features until all missing values are imputed or the imputation cost is exhausted. Experimental results on UCI data sets demonstrated the advantages of our proposed DIM, compared to the comparison methods, in terms of prediction accuracy and classification accuracy.
DOI: 10.1109/ijcnn.2007.4371139
发表时间: 2007-10
期刊: 2007 International Joint Conference on Neural Networks
影响因子: --
作者:
Mostafa M. Hassan;A. Atiya;N. E. Gayar;R. El-Fouly
通讯作者: Mostafa M. Hassan;A. Atiya;N. E. Gayar;R. El-Fouly
在分类回归模型中,二进制/分类协变量的二进制/分类协变量的多重估算。
DOI: 10.1002/sim.8004
发表时间: 2019-02-28
影响因子: 2
作者:
Pham TM;Carpenter JR;Morris TP;Wood AM;Petersen I
通讯作者: Petersen I
DOI: 10.1109/jsyst.2016.2576026
发表时间: 2018-06
影响因子: 4.4
作者:
Liang Zhao;Zhikui Chen;Zhennan Yang;Yueming Hu;M. Obaidat
通讯作者: Liang Zhao;Zhikui Chen;Zhennan Yang;Yueming Hu;M. Obaidat
DOI: --
发表时间: 2002-12
期刊: ArXiv
影响因子: --
作者:
Peter D. Turney
通讯作者: Peter D. Turney
DOI: 10.1145/1390156.1390186
发表时间: 2008-07
期刊: --
影响因子: --
作者:
Uwe Dick;P. Haider;T. Scheffer
通讯作者: Uwe Dick;P. Haider;T. Scheffer