Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of Pessimism

Bridging Offline Reinforcement Learning and Imitation Learning: A Tale of Pessimism
复制标题

DOI:
10.1109/tit.2022.3185139
复制
发表时间:
2021-03
影响因子:
2.5
通讯作者:
Paria Rashidinejad;Banghua Zhu;Cong Ma;Jiantao Jiao;Stuart J. Russell
Paria Rashidinejad;Banghua Zhu;Cong Ma;Jiantao Jiao;Stuart J. Russell
中科院分区:
计算机科学2区
文献类型:
--
作者:
Paria Rashidinejad;Banghua Zhu;Cong Ma;Jiantao Jiao;Stuart J. Russell

文献摘要

被引文献

相似文献

离线增强学习(RL)算法试图从固定数据集中学习最佳策略,而无需主动数据收集。基于离线数据集的组成,使用了两种主要方法:适用于专家数据集的模仿学习,以及通常需要统一的覆盖范围数据集的香草离线RL。从实际的角度来看,数据集通常偏离这两个极端,而确切的数据组成通常是未知的。为了弥合这一差距,我们提出了一个新的离线RL框架,称为单极浓缩性,该框架在两个极端的数据组成之间平稳地插值,因此统一模仿学习和Vanilla Offline RL。在这个新的框架下,我们问:一个人可以开发一种算法,该算法可以适应未知数据组成的最低最佳速率?为了解决这个问题,我们认为在离线RL的不确定性面前根据悲观而开发的较低信心约束(LCB)算法。我们研究LCB的有限样本特性以及多军匪徒,上下文匪徒和马尔可夫决策过程(MDPS)中的信息理论限制。我们的分析揭示了有关最优率的令人惊讶的事实。特别是,在上下文的匪徒和RL中,LCB都达到了几乎专家数据集的快速收敛率,类似于通过模仿学习而获得的,与离线RL中达到的缓慢率相反。在上下文的强盗中,我们证明LCB对整个数据组成范围都是最佳的最佳选择,从而实现了从模仿学习到离线RL的平稳过渡。我们进一步表明,LCB在MDP中几乎是最佳的。
Offline reinforcement learning (RL) algorithms seek to learn an optimal policy from a fixed dataset without active data collection. Based on the composition of the offline dataset, two main methods are used: imitation learning which is suitable for expert datasets, and vanilla offline RL which often requires uniform coverage datasets. From a practical standpoint, datasets often deviate from these two extremes and the exact data composition is usually unknown. To bridge this gap, we present a new offline RL framework, called single-policy concentrability, that smoothly interpolates between the two extremes of data composition, hence unifying imitation learning and vanilla offline RL. Under this new framework, we ask: can one develop an algorithm that achieves a minimax optimal rate adaptive to unknown data composition? To address this question, we consider a lower confidence bound (LCB) algorithm developed based on pessimism in the face of uncertainty in offline RL. We study finite-sample properties of LCB as well as information-theoretic limits in multi-armed bandits, contextual bandits, and Markov decision processes (MDPs). Our analysis reveals surprising facts about optimality rates. In particular, in both contextual bandits and RL, LCB achieves a fast convergence rate for nearly-expert datasets, analogous to the one achieved by imitation learning, contrary to the slow rate achieved in offline RL. In contextual bandits, we prove that LCB is adaptively optimal for the entire data composition range, achieving a smooth transition from imitation learning to offline RL. We further show that LCB is almost adaptively optimal in MDPs.