Comparison of Random Forest and Parametric Imputation Models for Imputing Missing Data Using MICE: A CALIBER Study

Comparison of Random Forest and Parametric Imputation Models for Imputing Missing Data Using MICE: A CALIBER Study
复制标题

DOI:
10.1093/aje/kwt312
复制
发表时间:
2014-03-15
影响因子:
5
通讯作者:
Hemingway, Harry
Hemingway, Harry
中科院分区:
医学2区
文献类型:
--
作者:
Shah, Anoop D.;Bartlett, Jonathan W.;Hemingway, Harry

文献摘要

被引文献

相似文献

链式方程多变量插补(MICE)是流行病学研究中常用的缺失数据插补方法。“真实”插补模型可能包含默认插补模型中不包含的非线性。随机森林插补是一种机器学习技术,可以适应非线性和交互作用,并且不需要指定特定的回归模型。我们在2项模拟研究中比较了参数MICE与基于森林的随机MICE算法。第一项研究使用了从CALIBER数据库(使用关联定制研究和电子记录的心血管疾病研究; 2001-2010)中的10,128名稳定型心绞痛患者中抽取的2,000人的1,000份随机样本,所有协变量均具有完整的数据。人为地使变量“随机缺失”,并比较使用不同插补方法获得的参数估计值的偏倚和效率。这两种MICE方法都产生了无偏的(对数)风险比估计值,但随机森林更有效,产生的置信区间更窄。第二项研究使用了模拟数据,其中部分观测变量以非线性方式依赖于完全观测变量。使用随机森林MICE,参数估计值偏差较小,置信区间覆盖率更好。这表明,随机森林插补可能是有用的插补复杂的流行病学数据集,其中一些患者有缺失的数据。
Multivariate imputation by chained equations (MICE) is commonly used for imputing missing data in epidemiologic research. The "true" imputation model may contain nonlinearities which are not included in default imputation models. Random forest imputation is a machine learning technique which can accommodate nonlinearities and interactions and does not require a particular regression model to be specified. We compared parametric MICE with a random forest-based MICE algorithm in 2 simulation studies. The first study used 1,000 random samples of 2,000 persons drawn from the 10,128 stable angina patients in the CALIBER database (Cardiovascular Disease Research using Linked Bespoke Studies and Electronic Records; 2001-2010) with complete data on all covariates. Variables were artificially made "missing at random," and the bias and efficiency of parameter estimates obtained using different imputation methods were compared. Both MICE methods produced unbiased estimates of (log) hazard ratios, but random forest was more efficient and produced narrower confidence intervals. The second study used simulated data in which the partially observed variable depended on the fully observed variables in a nonlinear way. Parameter estimates were less biased using random forest MICE, and confidence interval coverage was better. This suggests that random forest imputation may be useful for imputing complex epidemiologic data sets in which some patients have missing data.