Multiple Imputation for Incomplete Data in Epidemiologic Studies

Multiple Imputation for Incomplete Data in Epidemiologic Studies
复制标题

DOI:
10.1093/aje/kwx349
复制
发表时间:
2018-03-01
影响因子:
5
通讯作者:
Schisterman, Enrique F.
Schisterman, Enrique F.
中科院分区:
医学2区
文献类型:
--
作者:
Harel, Ofer;Mitchell, Emily M.;Schisterman, Enrique F.

文献摘要

被引文献

相似文献

流行病学研究经常容易出现信息缺失。在流行病学研究中,忽略缺失变量的观察仍然是一种常见的策略,但如果值不是完全随机缺失的,这种简单的方法通常会严重偏倚感兴趣的参数估计值。即使当缺失是完全随机的,完整的情况下分析可以降低估计参数的效率,因为大量的可用数据简单地抛出与不完整的观察。减轻缺失信息影响的替代方法,如多重插补,正在成为一种越来越流行的策略,以保留所有可用的信息,减少潜在的偏差,并提高参数估计的效率。在本文中,我们描述了多重插补的理论基础,我们说明了这种方法的应用程序的一部分,合作的挑战,以评估性能的各种技术处理缺失数据(美国流行病学杂志。2018;187(3):568-575)。我们详细介绍了对围产期合作项目(1959-1974)的一部分数据进行多重插补所需的步骤,该项目的目标是估计与妊娠期间吸烟相关的自然流产的几率。
Epidemiologic studies are frequently susceptible to missing information. Omitting observations with missing variables remains a common strategy in epidemiologic studies, yet this simple approach can often severely bias parameter estimates of interest if the values are not missing completely at random. Even when missingness is completely random, complete-case analysis can reduce the efficiency of estimated parameters, because large amounts of available data are simply tossed out with the incomplete observations. Alternative methods for mitigating the influence of missing information, such as multiple imputation, are becoming an increasing popular strategy in order to retain all available information, reduce potential bias, and improve efficiency in parameter estimation. In this paper, we describe the theoretical underpinnings of multiple imputation, and we illustrate application of this method as part of a collaborative challenge to assess the performance of various techniques for dealing with missing data (Am J Epidemiol. 2018;187(3):568-575). We detail the steps necessary to perform multiple imputation on a subset of data from the Collaborative Perinatal Project (1959-1974), where the goal is to estimate the odds of spontaneous abortion associated with smoking during pregnancy.