Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning

Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning
复制标题

DOI:
10.1287/opre.2021.2249
复制
发表时间:
2019-09
期刊:
Oper. Res.
影响因子:
--
通讯作者:
Nathan Kallus;Masatoshi Uehara
Nathan Kallus;Masatoshi Uehara
中科院分区:
其他
文献类型:
--
作者:
Nathan Kallus;Masatoshi Uehara

文献摘要

被引文献

相似文献

在离线强化学习(RL)中,我们使用现有的非策略数据来评估和学习新策略,这在实验具有挑战性且模拟不可靠的应用中至关重要,例如医学。这也是众所周知的困难,因为观察到的轨迹和任何新政策产生的轨迹之间的相似性(密度比)随着地平线的增长呈指数级下降,称为地平线的诅咒,这严重限制了离线RL的应用,无论地平线是中等长度甚至无限长。在“用双重强化学习有效打破非政策评估中的地平线诅咒”中,Kallus和Uehara开始理解这些限制以及何时可以打破它们。他们通过推导不同模型中策略值估计问题的半参数效率下界来精确地描述诅咒。一方面,这表明了为什么诅咒必然困扰着标准估计:它们甚至在非马尔可夫模型中也有效,因此必须受到相应界限的限制。另一方面,在某些马尔可夫模型中,更高的效率是可能的,并且它们给出了在无限时域马尔可夫决策过程中实现这些低得多的效率界的第一估计。
Demystifying the Curse of Horizon in Offline Reinforcement Learning in Order to Break It Offline reinforcement learning (RL), where we evaluate and learn new policies using existing off-policy data, is crucial in applications where experimentation is challenging and simulation unreliable, such as medicine. It is also notoriously difficult because the similarity (density ratio) between observed trajectories and those generated by any new policy diminishes exponentially as the horizon grows, known as the curse of horizon, which severely limits the application of offline RL whenever horizons are moderately long or even infinite. In “Efficiently Breaking the Curse of Horizon in Off-Policy Evaluation with Double Reinforcement Learning,” Kallus and Uehara set out to understand these limits and when they can be broken. They precisely characterize the curse by deriving the semiparametric efficiency lower bounds for the policy-value estimation problem in different models. On the one hand, this shows why the curse necessarily plagues standard estimators: they work even in non-Markov models and therefore must be limited by the corresponding bound. On the other hand, greater efficiency is possible in certain Markovian models, and they give the first estimator achieving these much lower efficiency bounds in infinite-horizon Markov decision processes.