Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
复制标题

DOI:
10.48550/arxiv.2210.06718
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yuda Song;Yi Zhou;Ayush Sekhari;J. Bagnell;A. Krishnamurthy;Wen Sun
Yuda Song;Yi Zhou;Ayush Sekhari;J. Bagnell;A. Krishnamurthy;Wen Sun
中科院分区:
其他
文献类型:
--
作者:
Yuda Song;Yi Zhou;Ayush Sekhari;J. Bagnell;A. Krishnamurthy;Wen Sun

文献摘要

相似文献

我们考虑一种混合强化学习设置(Hybrid RL),其中代理可以访问离线数据集,并能够通过真实世界的在线交互收集经验。该框架缓解了纯离线和在线RL设置中出现的挑战,允许在理论和实践中设计简单且高效的算法。我们通过将经典的Q学习/迭代算法适应于混合设置来证明这些优势,我们称之为混合Q学习或Hy-Q。在我们的理论结果中,我们证明了该算法在计算和统计上都是有效的,只要离线数据集支持高质量的政策和环境有界双线性秩。值得注意的是,我们不需要对初始分布提供的覆盖率进行假设,这与策略梯度/迭代方法的保证相反。在我们的实验结果中,我们表明,具有神经网络函数逼近的Hy-Q在具有挑战性的基准测试(包括Montezuma的Revenge)上的性能优于最先进的在线、离线和混合RL基线。
We consider a hybrid reinforcement learning setting (Hybrid RL), in which an agent has access to an offline dataset and the ability to collect experience via real-world online interaction. The framework mitigates the challenges that arise in both pure offline and online RL settings, allowing for the design of simple and highly effective algorithms, in both theory and practice. We demonstrate these advantages by adapting the classical Q learning/iteration algorithm to the hybrid setting, which we call Hybrid Q-Learning or Hy-Q. In our theoretical results, we prove that the algorithm is both computationally and statistically efficient whenever the offline dataset supports a high-quality policy and the environment has bounded bilinear rank. Notably, we require no assumptions on the coverage provided by the initial distribution, in contrast with guarantees for policy gradient/iteration methods. In our experimental results, we show that Hy-Q with neural network function approximation outperforms state-of-the-art online, offline, and hybrid RL baselines on challenging benchmarks, including Montezuma's Revenge.