Offline Reinforcement Learning Under Value and Density-Ratio Realizability: the Power of Gaps

Offline Reinforcement Learning Under Value and Density-Ratio Realizability: the Power of Gaps
复制标题

DOI:
10.48550/arxiv.2203.13935
复制
发表时间:
2022-03
期刊:
--
影响因子:
--
通讯作者:
Jinglin Chen;Nan Jiang
Jinglin Chen;Nan Jiang
中科院分区:
其他
文献类型:
--
作者:
Jinglin Chen;Nan Jiang

文献摘要

相似文献

我们考虑了离线强化学习(RL)中一个具有挑战性的理论问题:在函数逼近器的可实现性假设下,在缺乏足够覆盖的数据集上获得样本效率保证。虽然现有的理论已经解决了学习下的可实现性和非探索性数据分别,没有工作已经能够同时解决这两个问题(除了并发工作,我们详细比较)。在附加间隙假设下,基于边缘化重要性抽样(MIS)形成的版本空间,对一个简单的悲观算法提供了保证,该保证只要求数据覆盖最优策略,函数类实现最优值和密度比函数.虽然类似的间隙假设已被用于RL理论的其他领域,我们的工作是第一个确定的实用程序和离线RL与弱函数近似间隙假设的新机制。
We consider a challenging theoretical problem in offline reinforcement learning (RL): obtaining sample-efficiency guarantees with a dataset lacking sufficient coverage, under only realizability-type assumptions for the function approximators. While the existing theory has addressed learning under realizability and under non-exploratory data separately, no work has been able to address both simultaneously (except for a concurrent work which we compare in detail). Under an additional gap assumption, we provide guarantees to a simple pessimistic algorithm based on a version space formed by marginalized importance sampling (MIS), and the guarantee only requires the data to cover the optimal policy and the function classes to realize the optimal value and density-ratio functions. While similar gap assumptions have been used in other areas of RL theory, our work is the first to identify the utility and the novel mechanism of gap assumptions in offline RL with weak function approximation.