Covariate selection for the nonparametric estimation of an average treatment effect

Covariate selection for the nonparametric estimation of an average treatment effect
复制标题

DOI:
10.1093/biomet/asr041
复制
发表时间:
2011-12-01
期刊:
影响因子:
2.7
通讯作者:
Richardson, Thomas S.
Richardson, Thomas S.
中科院分区:
数学2区
文献类型:
--
作者:
De Luna, Xavier;Waernbaum, Ingeborg;Richardson, Thomas S.

文献摘要

被引文献

相似文献

估计非随机治疗对感兴趣结果的影响的观察性研究在劳动经济学和流行病学等领域很常见。在控制给定的一组观察到的治疗前协变量时,此类研究通常依赖于无混杂治疗的假设。为了保证无混杂性而选择要控制的协变量应主要基于主题理论,尽管后者通常只提供部分指导。人们很容易在控制集中包含许多协变量,以尝试使无混杂治疗的假设变得现实。当非参数估计二元处理的效果时,包括不必要的协变量是次优的。例如,当使用 n(1/2) 一致估计量时,使用与无混杂性假设无关的协变量可能会导致效率损失。此外,当使用许多协变量时,偏差可能会主导方差。采用通常与治疗效果的非参数估计量结合使用的 Neyman-Rubin 模型,我们描述了原始协变量库中的子集,这些子集是最小的,因为在给定这些最小集合的任何适当子集的情况下,治疗不再是无混杂的。这些协变量子集被证明是在温和的假设下识别的。这些结果促使我们提出数据驱动算法来选择最小协变量集。
Observational studies in which the effect of a nonrandomized treatment on an outcome of interest is estimated are common in domains such as labour economics and epidemiology. Such studies often rely on an assumption of unconfounded treatment when controlling for a given set of observed pre-treatment covariates. The choice of covariates to control in order to guarantee unconfoundedness should primarily be based on subject matter theories, although the latter typically give only partial guidance. It is tempting to include many covariates in the controlling set to try to make the assumption of an unconfounded treatment realistic. Including unnecessary covariates is suboptimal when the effect of a binary treatment is estimated nonparametrically. For instance, when using a n(1/2)-consistent estimator, a loss of efficiency may result from using covariates that are irrelevant for the unconfoundedness assumption. Moreover, bias may dominate the variance when many covariates are used. Embracing the Neyman-Rubin model typically used in conjunction with nonparametric estimators of treatment effects, we characterize subsets from the original reservoir of covariates that are minimal in the sense that the treatment ceases to be unconfounded given any proper subset of these minimal sets. These subsets of covariates are shown to be identified under mild assumptions. These results lead us to propose data-driven algorithms for the selection of minimal sets of covariates.