Oracle Inequalities for Model Selection in Offline Reinforcement Learning

Oracle Inequalities for Model Selection in Offline Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2211.02016
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Jonathan Lee;G. Tucker;Ofir Nachum;Bo Dai;E. Brunskill
Jonathan Lee;G. Tucker;Ofir Nachum;Bo Dai;E. Brunskill
中科院分区:
其他
文献类型:
--
作者:
Jonathan Lee;G. Tucker;Ofir Nachum;Bo Dai;E. Brunskill

文献摘要

相似文献

在离线强化学习(RL)中,学习者利用先前记录的数据来学习良好的策略,而无需与环境交互。在实践中应用此类方法的一个主要挑战是缺乏用于模型选择和评估的理论原理和实践工具。为了解决这个问题,我们研究了价值函数逼近的离线强化学习中的模型选择问题。学习者被给予模型类的嵌套序列,以最小化贝尔曼平方误差,并且必须在其中进行选择以实现类的近似误差和估计误差之间的平衡。我们提出了第一个用于离线强化学习的模型选择算法,该算法实现了达到对数因子的极小最大速率最优预言不等式。 ModBE 算法将候选模型类的集合和通用基础离线 RL 算法作为输入。通过使用一种新颖的单方面泛化测试来连续消除模型类,ModBE 返回了一种策略,该策略具有随最小完整模型类的复杂性缩放的遗憾。除了理论保证之外,它在概念上简单且计算高效,相当于解决一系列平方损失回归问题,然后比较类之间的相对平方损失。我们通过几次数值模拟得出结论,表明它能够可靠地选择一个好的模型类。
In offline reinforcement learning (RL), a learner leverages prior logged data to learn a good policy without interacting with the environment. A major challenge in applying such methods in practice is the lack of both theoretically principled and practical tools for model selection and evaluation. To address this, we study the problem of model selection in offline RL with value function approximation. The learner is given a nested sequence of model classes to minimize squared Bellman error and must select among these to achieve a balance between approximation and estimation error of the classes. We propose the first model selection algorithm for offline RL that achieves minimax rate-optimal oracle inequalities up to logarithmic factors. The algorithm, ModBE, takes as input a collection of candidate model classes and a generic base offline RL algorithm. By successively eliminating model classes using a novel one-sided generalization test, ModBE returns a policy with regret scaling with the complexity of the minimally complete model class. In addition to its theoretical guarantees, it is conceptually simple and computationally efficient, amounting to solving a series of square loss regression problems and then comparing relative square loss between classes. We conclude with several numerical simulations showing it is capable of reliably selecting a good model class.