Missing data imputation using statistical and machine learning methods in a real breast cancer problem

Missing data imputation using statistical and machine learning methods in a real breast cancer problem
复制标题

DOI:
10.1016/j.artmed.2010.05.002
复制
发表时间:
2010-10-01
影响因子:
7.5
通讯作者:
Franco, Leonardo
Franco, Leonardo
中科院分区:
工程技术1区
文献类型:
--
作者:
Jerez, Jose M.;Molina, Ignacio;Franco, Leonardo

文献摘要

被引文献

相似文献

目的:缺失数据插补是一项重要任务,在这种情况下,必须使用所有可用的数据,而不是丢弃缺失值的记录。这项工作评估了几种统计和机器学习插补方法的性能,这些方法用于预测广泛的真实的乳腺癌数据集中患者的复发。基于统计技术的估算方法,例如,通过“El Alamo-I”项目收集的数据应用了平均值、热甲板和多重插补以及机器学习技术,例如多层感知器(MLP)、自组织映射(SOM)和k-最近邻(KNN),然后将结果与从列表删除(LD)插补方法获得的结果进行比较。该数据库包括来自3679名妇女的人口统计学,治疗和复发生存信息,这些妇女在属于西班牙乳腺癌研究小组(GEICAM)的32家不同医院诊断为可手术浸润性乳腺癌。结果:基于机器学习算法的插补方法在预测患者预后方面优于插补统计方法。弗里德曼的检验显示,(p = 0.0091),成对比较检验显示MLP、KNN和SOM的AUC显著高于(分别为p = 00053、p = 00048和p = 00071)。基于机器学习技术的方法最适合缺失值的插补,与基于统计程序的插补方法相比,可以显著提高预后准确性。(C)2010爱思唯尔有限公司版权所有。
Objectives: Missing data imputation is an important task in cases where it is crucial to use all available data and not discard records with missing values. This work evaluates the performance of several statistical and machine learning imputation methods that were used to predict recurrence in patients in an extensive real breast cancer data setMaterials and methods. Imputation methods based on statistical techniques, e.g., mean, hot-deck and multiple imputation, and machine learning techniques, e.g. multi-layer perceptron (MLP), self-organisation maps (SOM) and k-nearest neighbour (KNN), were applied to data collected through the "El Alamo-I" project, and the results were then compared to those obtained from the listwise deletion (LD) imputation method. The database includes demographic, therapeutic and recurrence-survival information from 3679 women with operable invasive breast cancer diagnosed in 32 different hospitals belonging to the Spanish Breast Cancer Research Group (GEICAM). The accuracies of predictions on early cancer relapse were measured using artificial neural networks (ANNs), in which different ANNs were estimated using the data sets with imputed missing valuesResults: The imputation methods based on machine learning algorithms outperformed imputation statistical methods in the prediction of patient outcome. Friedman's test revealed a significant difference (p = 0.0091) in the observed area under the ROC curve (AUC) values, and the pairwise comparison test showed that the AUCs for MLP, KNN and SOM were significantly higher (p = 00053, p = 0 0048 and p = 00071, respectively) than the AUC from the LD-based prognosis model.Conclusion: The methods based on machine learning techniques were the most suited for the imputation of missing values and led to a significant enhancement of prognosis accuracy compared to imputation methods based on statistical procedures. (C) 2010 Elsevier B.V. All rights reserved.