Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values

Bias and efficiency of multiple imputation compared with complete-case analysis for missing covariate values
复制标题

DOI:
10.1002/sim.3944
复制
发表时间:
2010-12-10
影响因子:
2
通讯作者:
Carlin, John B.
Carlin, John B.
中科院分区:
医学3区
文献类型:
--
作者:
White, Ian R.;Carlin, John B.

文献摘要

被引文献

相似文献

当回归模型中的一个或多个协变量存在缺失数据时,多重插补(MI)被广泛提倡作为对完全病例分析(CC)的改进。我们使用理论论证和模拟研究将这些方法与随机缺失情况下的MI进行比较。并且MI在广泛的场景中比CC更有效。对于其他缺失数据机制,在一种或两种方法中都会出现偏差。在我们的模拟设置中,当数据随机缺失时,CC偏向于空值。然而,当缺失与给定协变量的结果无关时,CC具有可忽略的偏倚,MI偏离零。在更一般的缺失数据机制下,MI的偏倚往往小于CC。由于MI在缺失协变量问题上并不总是优于CC,方法的选择应考虑在特定的实质性应用中已知的缺失数据机制。重要的是,方法的选择不应基于标准误差的比较。提出了新的方法来理解MI和CC之间的经验差异,这可能会提供对每种方法所基于的假设的适当性的见解,我们提出了一个新的指标,用于评估MI在精度方面可能获得的收益,即协变量(FICO)观测值中不完整病例的比例版权所有(C)2010 John Wiley & Sons,Ltd
When missing data occur in one or more covariates in a regression model multiple imputation (MI) is widely advocated as an improvement over complete case analysis (CC) We use theoretical arguments and simulation studies to compare these methods with MI implemented under a missing at random assumptionWhen data are missing completely at random both methods have negligible bias, and MI is more efficient than CC across a wide range of scenarios For other missing data mechanisms, bias arises in one or both methods In our simulation setting, CC is biased towards the null when data are missing at random However, when missingness is independent of the outcome given the covariates, CC has negligible bias and MI is biased away from the null With more general missing data mechanisms, bias tends to be smaller for MI than for CCSince MI is not always better than CC for missing covariate problems, the choice of method should take into account what is known about the missing data mechanism in a particular substantive application Importantly the choice of method should not be based on comparison of standard errors We propose new ways to understand empirical differences between MI and CC, which may provide insights into the appropriateness of the assumptions underlying each method, and we propose a new index for assessing the likely gain in precision from MI the fraction of Incomplete cases among the observed values of a covariate (FICO) Copyright (C) 2010 John Wiley & Sons, Ltd