Multiple Imputation for Missing Data via Sequential Regression Trees

Multiple Imputation for Missing Data via Sequential Regression Trees
复制标题

DOI:
10.1093/aje/kwq260
复制
发表时间:
2010-11-01
影响因子:
5
通讯作者:
Reiter, Jerome P.
Reiter, Jerome P.
中科院分区:
医学2区
文献类型:
--
作者:
Burgette, Lane F.;Reiter, Jerome P.

文献摘要

被引文献

相似文献

多重插补特别适合处理大型流行病学研究中的缺失数据,因为这些研究通常支持许多数据用户进行的广泛分析。其中一些分析可能涉及复杂的建模,包括相互作用和非线性关系。识别这种关系并将其编码到插补模型中,例如,在通过链式方程进行多重插补的条件回归中,对于大量的分类和连续变量来说,这可能是一项艰巨的任务。本文提出了一种利用序贯回归树作为条件模型,通过链式方程实现多重插补的非参数方法。这有可能捕获复杂的关系,数据输入者只需进行最小的调整。使用模拟,作者证明,该方法可以导致更合理的插补,因此更可靠的推论,在复杂的设置比天真的应用标准序贯回归插补技术。他们应用这种方法来估算100多个临床和调查变量的不良出生结果数据中的缺失值。他们使用后验预测检查和几个感兴趣的流行病学分析来评估插补。
Multiple imputation is particularly well suited to deal with missing data in large epidemiologic studies, because typically these studies support a wide range of analyses by many data users. Some of these analyses may involve complex modeling, including interactions and nonlinear relations. Identifying such relations and encoding them in imputation models, for example, in the conditional regressions for multiple imputation via chained equations, can be daunting tasks with large numbers of categorical and continuous variables. The authors present a nonparametric approach for implementing multiple imputation via chained equations by using sequential regression trees as the conditional models. This has the potential to capture complex relations with minimal tuning by the data imputer. Using simulations, the authors demonstrate that the method can result in more plausible imputations, and hence more reliable inferences, in complex settings than the naive application of standard sequential regression imputation techniques. They apply the approach to impute missing values in data on adverse birth outcomes with more than 100 clinical and survey variables. They evaluate the imputations using posterior predictive checks with several epidemiologic analyses of interest.