POLICY LEARNING WITH OBSERVATIONAL DATA

POLICY LEARNING WITH OBSERVATIONAL DATA
复制标题

DOI:
10.3982/ecta15732
复制
发表时间:
2021-01-01
期刊:
影响因子:
6.1
通讯作者:
Wager, Stefan
Wager, Stefan
中科院分区:
经济学1区
文献类型:
--
作者:
Athey, Susan;Wager, Stefan

文献摘要

被引文献

相似文献

在许多领域,从业者寻求使用观察数据来学习满足应用特定约束的治疗分配策略,例如预算、公平性、简单性或其他功能形式约束。例如,策略可能被限制为基于一组有限的易于观察的个体特征的决策树的形式。在半参数有效估计理论的激励下,我们提出了一种解决这一问题的新方法。我们的方法可用于优化二元处理或无穷小的连续处理,并且可以利用观察数据,其中使用各种策略确定因果关系,包括对可观测值和工具变量的选择。给定分配给每个人治疗的因果效应的双鲁棒估计量,我们开发了一个选择治疗对象的算法,并为最终政策的渐近功利后悔建立了强有力的保证。
In many areas, practitioners seek to use observational data to learn a treatment assignment policy that satisfies application-specific constraints, such as budget, fairness, simplicity, or other functional form constraints. For example, policies may be restricted to take the form of decision trees based on a limited set of easily observable individual characteristics. We propose a new approach to this problem motivated by the theory of semiparametrically efficient estimation. Our method can be used to optimize either binary treatments or infinitesimal nudges to continuous treatments, and can leverage observational data where causal effects are identified using a variety of strategies, including selection on observables and instrumental variables. Given a doubly robust estimator of the causal effect of assigning everyone to treatment, we develop an algorithm for choosing whom to treat, and establish strong guarantees for the asymptotic utilitarian regret of the resulting policy.