Recovery of information from multiple imputation: a simulation study.

Recovery of information from multiple imputation: a simulation study.
复制标题

DOI:
10.1186/1742-7622-9-3
复制
发表时间:
2012-06-13
影响因子:
2.3
通讯作者:
Carlin JB
Carlin JB
中科院分区:
其他
文献类型:
--
作者:
Lee KJ;Carlin JB

文献摘要

被引文献

相似文献

多重插补在处理缺失数据方面越来越受欢迎。然而,它的实施往往没有充分考虑它是否提供了任何优势,比完整的案例分析的研究问题的兴趣,或潜在的收益是否可能被抵消的偏差,从一个拟合不佳的插补模型,特别是随着缺失数据的数量增加。使用从合成人群中提取的模拟数据集(n = 1000),探索当在高度偏斜的连续协变量或二进制暴露中随机缺失不同比例的数据(10-90%)时,在估计二进制暴露变量的系数时,多重插补的信息恢复。使用多变量正态插补(MVNI)进行插补,采用简单或零偏对数转换来管理非正态性。在多重插补和完整病例分析之间比较了一组回归参数估计值的偏倚、精密度、均方误差和覆盖率。对于连续协变量中的缺失,与完整病例分析相比,多重插补对二元暴露变量的影响产生的偏倚更小,精度更高,缺失数据越多,精度提高越多。然而,即使只有中度缺失,在估计连续协变量的影响时,当偏度没有得到充分解决时,大的偏倚和大量的覆盖不足是明显的。对于二进制协变量的缺失,所有估计值的偏倚可忽略不计,但多重插补的精密度增益极小,特别是对于二进制暴露的系数。虽然如果混杂调整所需的协变量缺失,多重插补可能有用,但当关注的暴露变量数据缺失时,获益可能很小。此外,当存在大量缺失时,多重插补可能变得不可靠,并且如果插补模型不合适,则会引入完整病例分析中不存在的偏倚。处理缺失数据的流行病学家应该记住多重插补的潜在局限性和潜在益处。需要进一步开展工作,为有效应用这一方法提供更明确的准则。
Multiple imputation is becoming increasingly popular for handling missing data. However, it is often implemented without adequate consideration of whether it offers any advantage over complete case analysis for the research question of interest, or whether potential gains may be offset by bias from a poorly fitting imputation model, particularly as the amount of missing data increases. Simulated datasets (n = 1000) drawn from a synthetic population were used to explore information recovery from multiple imputation in estimating the coefficient of a binary exposure variable when various proportions of data (10-90%) were set missing at random in a highly-skewed continuous covariate or in the binary exposure. Imputation was performed using multivariate normal imputation (MVNI), with a simple or zero-skewness log transformation to manage non-normality. Bias, precision, mean-squared error and coverage for a set of regression parameter estimates were compared between multiple imputation and complete case analyses. For missingness in the continuous covariate, multiple imputation produced less bias and greater precision for the effect of the binary exposure variable, compared with complete case analysis, with larger gains in precision with more missing data. However, even with only moderate missingness, large bias and substantial under-coverage were apparent in estimating the continuous covariate’s effect when skewness was not adequately addressed. For missingness in the binary covariate, all estimates had negligible bias but gains in precision from multiple imputation were minimal, particularly for the coefficient of the binary exposure. Although multiple imputation can be useful if covariates required for confounding adjustment are missing, benefits are likely to be minimal when data are missing in the exposure variable of interest. Furthermore, when there are large amounts of missingness, multiple imputation can become unreliable and introduce bias not present in a complete case analysis if the imputation model is not appropriate. Epidemiologists dealing with missing data should keep in mind the potential limitations as well as the potential benefits of multiple imputation. Further work is needed to provide clearer guidelines on effective application of this method.