Bellman-consistent Pessimism for Offline Reinforcement Learning

Bellman-consistent Pessimism for Offline Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Tengyang Xie;Ching-An Cheng;Nan Jiang;Paul Mineiro;Alekh Agarwal
Tengyang Xie;Ching-An Cheng;Nan Jiang;Paul Mineiro;Alekh Agarwal
中科院分区:
其他
文献类型:
--
作者:
Tengyang Xie;Ching-An Cheng;Nan Jiang;Paul Mineiro;Alekh Agarwal

文献摘要

被引文献

相似文献

当对数据集缺乏详尽探索的推理时,最近在离线增强学习中变得突出。尽管它增加了算法,但过于悲观的推理在排除良好政策的发现方面可能同样损害,这对于流行的基于奖金的悲观主义而言是一个问题。在本文中,我们介绍了贝尔曼一致的悲观情绪的概念,以实现一般函数近似:而不是计算值的值函数的点下限,而是在与贝尔曼方程一致的函数集上实现了在初始状态的悲观情绪。我们的理论保证只需要在探索环境中作为标准钟表的封闭性,在这种情况下,基于奖励的悲观主义无法提供保证。即使在具有更强表达性假设的特殊情况下,我们的结果也可以通过$ \ Mathcal {o}(d)$在其样品复杂性中的最新基于奖励的方法来改善。值得注意的是,我们的算法会自动适应事后的最佳偏见差异权衡,而大多数先前的方法都需要先验调整额外的超标剂。
The use of pessimism, when reasoning about datasets lacking exhaustive exploration has recently gained prominence in offline reinforcement learning. Despite the robustness it adds to the algorithm, overly pessimistic reasoning can be equally damaging in precluding the discovery of good policies, which is an issue for the popular bonus-based pessimism. In this paper, we introduce the notion of Bellman-consistent pessimism for general function approximation: instead of calculating a point-wise lower bound for the value function, we implement pessimism at the initial state over the set of functions consistent with the Bellman equations. Our theoretical guarantees only require Bellman closedness as standard in the exploratory setting, in which case bonus-based pessimism fails to provide guarantees. Even in the special case of linear function approximation where stronger expressivity assumptions hold, our result improves upon a recent bonus-based approach by $\mathcal{O}(d)$ in its sample complexity when the action space is finite. Remarkably, our algorithms automatically adapt to the best bias-variance tradeoff in the hindsight, whereas most prior approaches require tuning extra hyperparameters a priori.