Bellman Residual Orthogonalization for Offline Reinforcement Learning

Bellman Residual Orthogonalization for Offline Reinforcement Learning
复制标题

DOI:
10.48550/arxiv.2203.12786
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
A. Zanette;M. Wainwright
A. Zanette;M. Wainwright
中科院分区:
其他
文献类型:
--
作者:
A. Zanette;M. Wainwright

文献摘要

相似文献

我们提出并分析了一个强化学习原理,该原理通过仅沿着用户定义的测试函数空间来执行Bellman方程的有效性来近似Bellman方程。专注于无模型离线RL与函数近似的应用程序,我们利用这一原则来获得置信区间的政策评估,以及在规定的政策类内的政策进行优化。我们证明了一个预言不等式我们的政策优化过程中的任意比较政策的价值和不确定性之间的权衡。测试函数空间的不同选择使我们能够在一个共同的框架内解决不同的问题。我们的特点的效率损失,从政策上的政策,关闭数据使用我们的程序,并建立连接到集中系数在过去的工作中研究。我们深入研究我们的方法与线性函数近似的实现,并提供理论保证多项式时间的实现,即使贝尔曼封闭不成立。
We propose and analyze a reinforcement learning principle that approximates the Bellman equations by enforcing their validity only along an user-defined space of test functions. Focusing on applications to model-free offline RL with function approximation, we exploit this principle to derive confidence intervals for off-policy evaluation, as well as to optimize over policies within a prescribed policy class. We prove an oracle inequality on our policy optimization procedure in terms of a trade-off between the value and uncertainty of an arbitrary comparator policy. Different choices of test function spaces allow us to tackle different problems within a common framework. We characterize the loss of efficiency in moving from on-policy to off-policy data using our procedures, and establish connections to concentrability coefficients studied in past work. We examine in depth the implementation of our methods with linear function approximation, and provide theoretical guarantees with polynomial-time implementations even when Bellman closure does not hold.