Efficient Evaluation of Natural Stochastic Policies in Offline Reinforcement Learning

Efficient Evaluation of Natural Stochastic Policies in Offline Reinforcement Learning
复制标题

DOI:
10.1093/biomet/asad059
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Nathan Kallus;Masatoshi Uehara
Nathan Kallus;Masatoshi Uehara
中科院分区:
其他
文献类型:
--
作者:
Nathan Kallus;Masatoshi Uehara

文献摘要

被引文献

相似文献

我们研究了自然随机策略的有效非策略评估,该策略定义为与未知行为策略的偏差。这与关于非政策评估的文献不同,后者主要考虑对明确指定的政策的评估。至关重要的是,使用自然随机策略的离线强化学习可以帮助缓解弱重叠问题,导致策略建立在当前实践的基础上,并提高策略在实践中的可实施性。与经典的预先指定的评估策略相比,在评估自然随机策略时,由于评估策略本身是未知的,衡量最佳估计误差的效率界被夸大了。本文给出了两类主要的自然随机策略:倾斜策略和修正治疗策略的有效界。然后,我们提出了有效的非参数估计量,它在宽松条件下达到了效率界,并且具有部分双重稳健性。
We study the efficient off-policy evaluation of natural stochastic policies, which are defined in terms of deviations from the unknown behaviour policy. This is a departure from the literature on off-policy evaluation that largely consider the evaluation of explicitly specified policies. Crucially, offline reinforcement learning with natural stochastic policies can help alleviate issues of weak overlap, lead to policies that build upon current practice, and improve policies' implementability in practice. Compared with the classic case of a prespecified evaluation policy, when evaluating natural stochastic policies, the efficiency bound, which measures the best-achievable estimation error, is inflated since the evaluation policy itself is unknown. In this paper we derive the efficiency bounds of two major types of natural stochastic policies: tilting policies and modified treatment policies. We then propose efficient nonparametric estimators that attain the efficiency bounds under lax conditions and enjoy a partial double robustness property.