Bayesian Reinforcement Learning in Factored POMDPs

Bayesian Reinforcement Learning in Factored POMDPs
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Sammie Katt;F. Oliehoek;Chris Amato
Sammie Katt;F. Oliehoek;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Sammie Katt;F. Oliehoek;Chris Amato

文献摘要

被引文献

相似文献

基于模型的贝叶斯强化学习(BRL)提供了处理勘探与开发权衡的原则性解决方案,但此类方法通常假设完全可观察的环境。少数适用于部分可观察域的贝叶斯RL方法,如贝叶斯自适应POMDP(BA-POMDP),规模很差。为了解决这个问题,我们引入了因子BA-POMDP模型(FBA-POMDP),这是一个框架,能够通过利用POMDP的底层结构来学习一个紧凑的动态模型。的FBA-POMDP框架铸造的问题作为一个规划任务,我们适应蒙特-卡罗树搜索规划算法,并开发了一个信念跟踪方法来近似的状态和模型变量的联合后验。我们的实证结果表明,该方法优于一些BRL基线,并且能够在因子分解已知时有效地学习,以及同时学习因子分解和模型参数。
Model-based Bayesian Reinforcement Learning (BRL) provides a principled solution to dealing with the exploration-exploitation trade-off, but such methods typically assume a fully observable environments. The few Bayesian RL methods that are applicable in partially observable domains, such as the Bayes-Adaptive POMDP (BA-POMDP), scale poorly. To address this issue, we introduce the Factored BA-POMDP model (FBA-POMDP), a framework that is able to learn a compact model of the dynamics by exploiting the underlying structure of a POMDP. The FBA-POMDP framework casts the problem as a planning task, for which we adapt the Monte-Carlo Tree Search planning algorithm and develop a belief tracking method to approximate the joint posterior over the state and model variables. Our empirical results show that this method outperforms a number of BRL baselines and is able to learn efficiently when the factorization is known, as well as learn both the factorization and the model parameters simultaneously.