Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning

Provable Benefits of Actor-Critic Methods for Offline Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
A. Zanette;M. Wainwright;E. Brunskill
A. Zanette;M. Wainwright;E. Brunskill
中科院分区:
其他
文献类型:
--
作者:
A. Zanette;M. Wainwright;E. Brunskill

文献摘要

相似文献

Actor-Critic方法广泛应用于离线强化学习实践中,但在理论上并没有得到很好的理解。我们提出了一种新的离线演员批评算法,它自然地结合了悲观主义原则,与现有技术相比具有几个关键优势。当贝尔曼评估算子相对于行动者策略的行动价值函数闭合时,该算法可以运行;这是比低阶 MDP 模型更通用的设置。尽管增加了通用性,但该过程在计算上是易于处理的,因为它涉及一系列二阶程序的求解。我们证明了该过程返回的策略的次优差距的上限,该上限取决于任何任意的、可能依赖于数据的比较器策略的数据覆盖范围。可实现的保证由与对数因子匹配的极小极大下界补充。
Actor-critic methods are widely used in offline reinforcement learning practice, but are not so well-understood theoretically. We propose a new offline actor-critic algorithm that naturally incorporates the pessimism principle, leading to several key advantages compared to the state of the art. The algorithm can operate when the Bellman evaluation operator is closed with respect to the action value function of the actor's policies; this is a more general setting than the low-rank MDP model. Despite the added generality, the procedure is computationally tractable as it involves the solution of a sequence of second-order programs. We prove an upper bound on the suboptimality gap of the policy returned by the procedure that depends on the data coverage of any arbitrary, possibly data dependent comparator policy. The achievable guarantee is complemented with a minimax lower bound that is matching up to logarithmic factors.