Variable Selection for Confounder Control, Flexible Modeling and Collaborative Targeted Minimum Loss-Based Estimation in Causal Inference.

Variable Selection for Confounder Control, Flexible Modeling and Collaborative Targeted Minimum Loss-Based Estimation in Causal Inference.
复制标题

DOI:
10.1515/ijb-2015-0017
复制
发表时间:
2016-05-01
期刊:
The international journal of biostatistics
影响因子:
--
通讯作者:
Gruber S
Gruber S
中科院分区:
其他
文献类型:
--
作者:
Schnitzer ME;Lok JJ;Gruber S

文献摘要

被引文献

相似文献

本文研究了在半参数模型中集成灵活倾向评分模型(非参数或机器学习方法)以估计因果量(例如治疗中的平均结果)的适当性。我们首先概述因果推理中基于知识和统计变量选择所涉及的一些问题,以及基于倾向得分拟合的自动选择的潜在陷阱。通过一个简单的例子,我们直接展示了使用治疗加权逆概率(IPTW)时调整暴露的纯粹原因的后果。当使用简单的方法选择倾向得分的模型时,可能会选择这些变量。我们描述了协作目标最小损失估计(C-TMLE)方法如何利用半参数有效估计量的协作双鲁棒性特性,根据条件结果模型中的误差选择倾向得分的协变量。最后,我们通过模拟研究比较了低维和高维设置中自动变量选择的几种方法。从这项模拟研究中,我们得出结论,使用 IPTW 对倾向评分进行灵活预测可能会导致估计效果较差,而基于目标最小损失的估计和 C-TMLE 可能会受益于灵活的预测,并且对于与治疗高度相关的变量的存在保持稳健。然而,在我们的研究中,基于标准影响函数的方差方法低估了标准误差,导致某些数据生成场景下的覆盖范围很差。
This paper investigates the appropriateness of the integration of flexible propensity score modeling (nonparametric or machine learning approaches) in semiparametric models for the estimation of a causal quantity, such as the mean outcome under treatment. We begin with an overview of some of the issues involved in knowledge-based and statistical variable selection in causal inference and the potential pitfalls of automated selection based on the fit of the propensity score. Using a simple example, we directly show the consequences of adjusting for pure causes of the exposure when using inverse probability of treatment weighting (IPTW). Such variables are likely to be selected when using a naive approach to model selection for the propensity score. We describe how the method of Collaborative Targeted minimum loss-based estimation (C-TMLE) capitalizes on the collaborative double robustness property of semiparametric efficient estimators to select covariates for the propensity score based on the error in the conditional outcome model. Finally, we compare several approaches to automated variable selection in low-and high-dimensional settings through a simulation study. From this simulation study, we conclude that using IPTW with flexible prediction for the propensity score can result in inferior estimation, while Targeted minimum loss-based estimation and C-TMLE may benefit from flexible prediction and remain robust to the presence of variables that are highly correlated with treatment. However, in our study, standard influence function-based methods for the variance underestimated the standard errors, resulting in poor coverage under certain data-generating scenarios.