More Efficient Policy Learning via Optimal Retargeting

More Efficient Policy Learning via Optimal Retargeting
复制标题

DOI:
10.1080/01621459.2020.1788948
复制
发表时间:
2020-08-02
影响因子:
3.7
通讯作者:
Kallus, Nathan
Kallus, Nathan
中科院分区:
数学1区
文献类型:
--
作者:
Kallus, Nathan

文献摘要

被引文献

相似文献

政策学习可用于从医疗保健、公民、电子商务等领域的观察数据中提取个性化的治疗方案。政策学习的一大障碍是不同行动的数据普遍缺乏重叠,这可能导致笨拙的政策评估和学习后的政策表现不佳。我们研究了一种基于重新定位的解决方案,即改变政策优化的人群。我们首先认为,在人口水平上,重新定位可能会导致很少或没有偏见。然后,我们描述了二元动作和多动作设置下的最优参考策略和重定向权重。我们根据新学习目标的渐近有效估计方差来做到这一点。我们进一步考虑额外控制重定向引起的潜在偏差的权重。一项模拟研究和一项个性化工作咨询案例研究的广泛实证结果表明,重新定位是一种相当简单的方法,可以显著改善应用于观察数据的任何政策学习过程。这篇文章可以在网上找到。
Policy learning can be used to extract individualized treatment regimes from observational data in healthcare, civics, e-commerce, and beyond. One big hurdle to policy learning is a commonplace lack of overlap in the data for different actions, which can lead to unwieldy policy evaluation and poorly performing learned policies. We study a solution to this problem based on retargeting, that is, changing the population on which policies are optimized. We first argue that at the population level, retargeting may induce little to no bias. We then characterize the optimal reference policy and retargeting weights in both binary-action and multi-action settings. We do this in terms of the asymptotic efficient estimation variance of the new learning objective. We further consider weights that additionally control for potential bias due to retargeting. Extensive empirical results in a simulation study and a case study of personalized job counseling demonstrate that retargeting is a fairly easy way to significantly improve any policy learning procedure applied to observational data.for this article are available online.