How many imputations are really needed? - Some practical clarifications of multiple imputation theory

How many imputations are really needed? - Some practical clarifications of multiple imputation theory
复制标题

DOI:
10.1007/s11121-007-0070-9
复制
发表时间:
2007-09-01
期刊:
影响因子:
3.5
通讯作者:
Gilreath, Tamika D.
Gilreath, Tamika D.
中科院分区:
医学2区
文献类型:
--
作者:
Graham, John W.;Olchowski, Allison E.;Gilreath, Tamika D.

文献摘要

被引文献

相似文献

多重填补(MI)和完全信息极大似然估计(FIML)是缺失数据分析的两种最常见方法。从理论上讲,当使用相同变量检验相同模型,且使用MI进行填补的次数m趋近于无穷大时,MI和FIML是等效的。然而,重要的是要了解在对预防科学家而言重要的方面MI和FIML足够等效之前需要进行多少次填补。MI理论表明,较小的m值,即使是3到5次填补,也能产生极好的结果。先前关于足够的m的指导原则是基于相对效率,它涉及被估计参数的缺失信息比例(γ)以及m。在本研究中,我们使用蒙特卡罗模拟在γ和m变化的几种情景下测试MI模型。感兴趣的回归系数的标准误和p值随m变化,但变化速率与相对效率不同。最重要的是,随着m变小,小效应量的统计功效降低,且这种功效下降的速率比相对效率变化所预测的要大得多。基于我们的研究结果,我们建议使用MI的研究人员应进行比先前认为足够的次数多得多的填补。这些建议基于γ,并考虑到由于填补次数过少而对可预防的功效下降(与FIML相比)的容忍度。
Multiple imputation (MI) and full information maximum likelihood (FIML) are the two most common approaches to missing data analysis. In theory, MI and FIML are equivalent when identical models are tested using the same variables, and when m, the number of imputations performed with MI, approaches infinity. However, it is important to know how many imputations are necessary before MI and FIML are sufficiently equivalent in ways that are important to prevention scientists. MI theory suggests that small values of m, even on the order of three to five imputations, yield excellent results. Previous guidelines for sufficient m are based on relative efficiency, which involves the fraction of missing information (gamma) for the parameter being estimated, and m. In the present study, we used a Monte Carlo simulation to test MI models across several scenarios in which gamma and m were varied. Standard errors and p-values for the regression coefficient of interest varied as a function of m, but not at the same rate as relative efficiency. Most importantly, statistical power for small effect sizes diminished as m became smaller, and the rate of this power falloff was much greater than predicted by changes in relative efficiency. Based our findings, we recommend that researchers using MI should perform many more imputations than previously considered sufficient. These recommendations are based on gamma, and take into consideration one's tolerance for a preventable power falloff (compared to FIML) due to using too few imputations.