Recursive partitioning for heterogeneous causal effects

Recursive partitioning for heterogeneous causal effects
复制标题

DOI:
10.1073/pnas.1510489113
复制
发表时间:
2016-07-05
影响因子:
11.1
通讯作者:
Imbens, Guido
Imbens, Guido
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Athey, Susan;Imbens, Guido

文献摘要

被引文献

相似文献

在本文中,我们提出了在实验性和观察性研究中估计因果效应异质性的方法,以及对总体子集间治疗效应差异大小进行假设检验的方法。我们提供了一种数据驱动的方法,将数据划分为治疗效应大小不同的亚群。该方法能够构建治疗效应的有效置信区间,即使相对于样本量存在许多协变量,且无需“稀疏性”假设。我们提出了一种“诚实”的估计方法,即使用一个样本构建划分,另一个样本估计每个亚群的治疗效应。我们的方法基于回归树方法,并进行了修改,以优化治疗效应的拟合优度并考虑诚实估计。我们的模型选择标准预期通过诚实估计消除偏差,并考虑在每个亚群内对治疗效应估计方差进行额外划分的影响。我们解决了这样一个挑战:对于任何单个单位都无法观察到因果效应的“基本事实”,因此必须修改交叉验证的标准方法。通过一项模拟研究,我们表明,对于我们首选的方法,诚实估计导致90%置信区间具有名义覆盖率,而对于非诚实方法,覆盖率在74%到84%之间。诚实估计需要用较小的样本量估计模型;我们首选方法在治疗效应均方误差方面的代价在7% - 22%之间。
In this paper we propose methods for estimating heterogeneity in causal effects in experimental and observational studies and for conducting hypothesis tests about the magnitude of differences in treatment effects across subsets of the population. We provide a data-driven approach to partition the data into subpopulations that differ in the magnitude of their treatment effects. The approach enables the construction of valid confidence intervals for treatment effects, even with many covariates relative to the sample size, and without "sparsity" assumptions. We propose an "honest" approach to estimation, whereby one sample is used to construct the partition and another to estimate treatment effects for each subpopulation. Our approach builds on regression tree methods, modified to optimize for goodness of fit in treatment effects and to account for honest estimation. Our model selection criterion anticipates that bias will be eliminated by honest estimation and also accounts for the effect of making additional splits on the variance of treatment effect estimates within each subpopulation. We address the challenge that the "ground truth" for a causal effect is not observed for any individual unit, so that standard approaches to cross-validation must be modified. Through a simulation study, we show that for our preferred method honest estimation results in nominal coverage for 90% confidence intervals, whereas coverage ranges between 74% and 84% for nonhonest approaches. Honest estimation requires estimating the model with a smaller sample size; the cost in terms of mean squared error of treatment effects for our preferred method ranges between 7-22%.