Tuning multiple imputation by predictive mean matching and local residual draws

Tuning multiple imputation by predictive mean matching and local residual draws
复制标题

DOI:
10.1186/1471-2288-14-75
复制
发表时间:
2014-06-05
影响因子:
4
通讯作者:
Royston, Patrick
Royston, Patrick
中科院分区:
医学3区
文献类型:
--
作者:
Morris, Tim P.;White, Ian R.;Royston, Patrick

文献摘要

被引文献

相似文献

背景资料:多重插补是处理不完全协变量的常用方法,因为它可以在数据随机缺失时提供有效的推断。这取决于能够正确指定用于估算缺失值的参数模型,这在许多现实环境中可能很困难。通过预测均值匹配(PMM)进行插补借用了具有相似预测均值的供体的观察值;通过局部残差抽取(LRD)进行插补则借用了供体的残差。这两种方法放宽了一些假设的参数插补,承诺更大的鲁棒性时,插补模型是misspecified.Methods:我们回顾发展的PMM和LRD和概述了各种形式,并旨在澄清一些选择如何以及何时应该使用它们。我们比较性能完全参数插补在模拟studies中,首先当插补模型是正确指定的,然后当它是misspecified.Results:在使用PMM或LRD,我们强烈警告不要使用一个单一的捐助者,在某些实现中的默认值,而不是提倡从一个池中的采样约10个捐助者。我们还澄清了哪个匹配度量是最好的。在目前的MI软件有几个穷人implementation.Conclusions:PMM和LRD可能有一个作用,插补协变量(i)这是不是与结果密切相关,(ii)当插补模型被认为是轻微的,但不是严重错误指定。研究人员应该努力正确地指定插补模型,而不是期望预测均值匹配或局部残差绘制来完成这项工作。
Background: Multiple imputation is a commonly used method for handling incomplete covariates as it can provide valid inference when data are missing at random. This depends on being able to correctly specify the parametric model used to impute missing values, which may be difficult in many realistic settings. Imputation by predictive mean matching (PMM) borrows an observed value from a donor with a similar predictive mean; imputation by local residual draws (LRD) instead borrows the donor's residual. Both methods relax some assumptions of parametric imputation, promising greater robustness when the imputation model is misspecified.Methods: We review development of PMM and LRD and outline the various forms available, and aim to clarify some choices about how and when they should be used. We compare performance to fully parametric imputation in simulation studies, first when the imputation model is correctly specified and then when it is misspecified.Results: In using PMM or LRD we strongly caution against using a single donor, the default value in some implementations, and instead advocate sampling from a pool of around 10 donors. We also clarify which matching metric is best. Among the current MI software there are several poor implementations.Conclusions: PMM and LRD may have a role for imputing covariates (i) which are not strongly associated with outcome, and (ii) when the imputation model is thought to be slightly but not grossly misspecified. Researchers should spend efforts on specifying the imputation model correctly, rather than expecting predictive mean matching or local residual draws to do the work.