Conformal inference of counterfactuals and individual treatment effects

Conformal inference of counterfactuals and individual treatment effects
复制标题

反事实的保角推断与个体治疗效果

DOI:
10.1111/rssb.12445
复制
发表时间:
2021-10-07
影响因子:
5.8
通讯作者:
Candes, Emmanuel J.
Candes, Emmanuel J.
中科院分区:
数学1区
文献类型:
--
作者:
Lei, Lihua;Candes, Emmanuel J.

文献摘要

被引文献

相似文献

评估治疗效果异质性广泛告知治疗决策。目前,非常重视通过灵活的机器学习算法估计条件平均治疗效果。虽然这些方法在一致性和收敛速度方面具有一定的理论吸引力,但它们在不确定性量化方面通常表现不佳。这是令人不安的,因为评估风险对于在敏感和不确定的环境中做出可靠的决策至关重要。在这项工作中,我们提出了一个共形推理为基础的方法,可以产生可靠的区间估计的反事实和个人的治疗效果下的潜在的结果框架。对于完全随机化或分层随机化实验,无论未知的数据生成机制如何,区间都保证了有限样本的平均覆盖率。对于随机化实验的可验证的遵守和一般的观察性研究服从强可验证性假设,区间满足一个双重稳健的属性,其中规定如下:平均覆盖率是近似控制的,如果可以准确估计的倾向得分或潜在结果的条件分位数。对合成数据集和真实的数据集的数值研究表明,即使在简单的模型中,现有的方法也存在显著的覆盖不足。相比之下,我们的方法以合理的短间隔实现了所需的覆盖范围。
Evaluating treatment effect heterogeneity widely informs treatment decision making. At the moment, much emphasis is placed on the estimation of the conditional average treatment effect via flexible machine learning algorithms. While these methods enjoy some theoretical appeal in terms of consistency and convergence rates, they generally perform poorly in terms of uncertainty quantification. This is troubling since assessing risk is crucial for reliable decision-making in sensitive and uncertain environments. In this work, we propose a conformal inference-based approach that can produce reliable interval estimates for counterfactuals and individual treatment effects under the potential outcome framework. For completely randomized or stratified randomized experiments with perfect compliance, the intervals have guaranteed average coverage in finite samples regardless of the unknown data generating mechanism. For randomized experiments with ignorable compliance and general observational studies obeying the strong ignorability assumption, the intervals satisfy a doubly robust property which states the following: the average coverage is approximately controlled if either the propensity score or the conditional quantiles of potential outcomes can be estimated accurately. Numerical studies on both synthetic and real data sets empirically demonstrate that existing methods suffer from a significant coverage deficit even in simple models. In contrast, our methods achieve the desired coverage with reasonably short intervals.