Comparative methods for handling missing data in large databases

Comparative methods for handling missing data in large databases
复制标题

DOI:
10.1016/j.jvs.2013.05.008
复制
发表时间:
2013-11-01
影响因子:
4.3
通讯作者:
Nguyen, Louis L.
Nguyen, Louis L.
中科院分区:
医学2区
文献类型:
--
作者:
Henry, Antonia J.;Hevelone, Nathanael D.;Nguyen, Louis L.

文献摘要

被引文献

相似文献

目的:分析复杂的调查数据库是卫生服务研究人员的重要工具。丢失数据元素是具有挑战性的,因为“丢失”的原因是多因素的,特别是种族等分类变量。我们模拟了RACE的缺失数据,并分析了预测严重肢体缺血(CLI)患者主要截肢手术的5种方法的偏差。方法:从2003年至2007年全国住院患者样本(加权n=684,057)中选择具有完整观察数据的患者出院,其中包括下肢血管重建术或重大截肢和CLI。在考虑几种随机缺失数据方案的情况下,我们比较了五种缺失数据方法:完全案例分析法、用观测频率替代法、缺失指标变量法、多重补充法和加权估计方程式法。我们创建了100个模拟数据集,其中5%、15%或30%的受试者种族从整个数据集中缺失。通过比较来自每种方法的100个模拟数据集的平均估计回归系数(Beta(缺失))与来自完全观察的数据集的估计值(Beta(完整))来估计偏差,相对偏差计算为(Beta(完整)-Beta(缺失)/Beta(完整))x 100%。结果:我们的结果表明,重新加权的估计方程产生的偏差最小,而缺失的指标变量产生的偏差系数最大。完整的病例分析,用观察到的频率替换,以及多重归因导致了适度的偏差。敏感度分析表明,最优方法的选择取决于所遇到的缺失数据的数量和类型。结论:缺失数据是大型数据库研究中的一个重要分析课题。常用的缺失指示变量法会带来严重的偏差,应谨慎使用。我们给出了经验证据来指导缺失数据处理方法的选择。
Objective: Analysis of complex survey databases is an important tool for health services researchers. Missing data elements are challenging because the reasons for "missingness" are multifactorial, especially categorical variables such as race. We simulated missing data for race and analyzed the bias from five methods used in predicting major amputation in patients with critical limb ischemia (CLI).Methods: Patient discharges with fully observed data containing lower extremity revascularization or major amputation and CLI were selected from the 2003 to 2007 Nationwide Inpatient Sample, a complex survey database (weighted n = 684,057). Considering several random missing data schemes, we compared five missing data methods: complete case analysis, replacement with observed frequencies, missing indicator variable, multiple imputation, and reweighted estimating equations. We created 100 simulated data sets, with 5%, 15%, or 30% of subjects' race drawn to be missing from the full data set. Bias was estimated by comparing the estimated regression coefficients averaged over 100 simulated data sets (beta(miss)) from each method vs estimates from the fully observed data set (beta(full)), with relative bias calculated as (beta(full)-beta(miss)/beta(full)) x 100%.Results: Our results demonstrate that reweighted estimating equations produce the least biased and the missing indicator variable produces the most biased coefficients. Complete case analysis, replacement with observed frequencies, and multiple imputation resulted in moderate bias. Sensitivity analysis demonstrated the optimal method choice depends on the quantity and type of missing data encountered.Conclusions: Missing data are an important analytic topic in research with large databases. The commonly used missing indicator variable method introduces severe bias and should be used with caution. We present empiric evidence to guide method selection for handling missing data.