Model-based Lifelong Reinforcement Learning with Bayesian Exploration

Model-based Lifelong Reinforcement Learning with Bayesian Exploration
复制标题

DOI:
10.48550/arxiv.2210.11579
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Haotian Fu;Shangqun Yu;Michael S. Littman;G. Konidaris
Haotian Fu;Shangqun Yu;Michael S. Littman;G. Konidaris
中科院分区:
其他
文献类型:
--
作者:
Haotian Fu;Shangqun Yu;Michael S. Littman;G. Konidaris

文献摘要

相似文献

我们提出了一种基于模型的终身强化学习方法,该方法估计了一个分层贝叶斯后验估计,提取了不同任务之间共享的共同结构。学习的后验知识与基于样本的贝叶斯探索过程相结合,提高了一系列相关任务的样本学习效率。我们首先分析了在有限MDP环境下样本复杂度与后验概率初始化质量之间的关系。接下来,我们通过引入一种变分贝叶斯终身强化学习算法将该方法扩展到连续状态域,该算法可以与最近基于模型的深度RL方法相结合,并且表现出向后迁移。在几个具有挑战性的领域上的实验结果表明,我们的算法取得了比最先进的终身RL方法更好的前向和后向传输性能。
We propose a model-based lifelong reinforcement-learning approach that estimates a hierarchical Bayesian posterior distilling the common structure shared across different tasks. The learned posterior combined with a sample-based Bayesian exploration procedure increases the sample efficiency of learning across a family of related tasks. We first derive an analysis of the relationship between the sample complexity and the initialization quality of the posterior in the finite MDP setting. We next scale the approach to continuous-state domains by introducing a Variational Bayesian Lifelong Reinforcement Learning algorithm that can be combined with recent model-based deep RL methods, and that exhibits backward transfer. Experimental results on several challenging domains show that our algorithms achieve both better forward and backward transfer performance than state-of-the-art lifelong RL methods.