BADDr: Bayes-Adaptive Deep Dropout RL for POMDPs

BADDr: Bayes-Adaptive Deep Dropout RL for POMDPs
复制标题

DOI:
10.5555/3535850.3535932
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Sammie Katt;Hai V. Nguyen;F. Oliehoek;Chris Amato
Sammie Katt;Hai V. Nguyen;F. Oliehoek;Chris Amato
中科院分区:
其他
文献类型:
--
作者:
Sammie Katt;Hai V. Nguyen;F. Oliehoek;Chris Amato

文献摘要

被引文献

相似文献

虽然强化学习(RL)在可扩展性方面取得了很大进展,但探索和部分可观测性仍然是活跃的研究课题。相比之下,贝叶斯RL(BRL)为状态估计和探索-开发权衡提供了原则性的答案,但难以扩展。为了应对这一挑战,已经提出了具有各种先前假设的BRL框架,并取得了不同的成功。这项工作提出了一个表示不可知的制定BRL下部分可观测性,统一以前的模型下一个理论的保护伞。为了证明其实际意义,我们还提出了一种新的推导,贝叶斯自适应深度丢弃rl(BADDr),基于丢弃网络。在这种参数化下,与以前的工作相比,对状态和动态的信念是一个更具可扩展性的推理问题。我们选择行动,通过蒙特-卡罗树搜索和经验表明,我们的方法是有竞争力的国家的最先进的BRL方法的小域,同时能够解决更大的。
While reinforcement learning (RL) has made great advances in scalability, exploration and partial observability are still active research topics. In contrast, Bayesian RL (BRL) provides a principled answer to both state estimation and the exploration-exploitation trade-off, but struggles to scale. To tackle this challenge, BRL frameworks with various prior assumptions have been proposed, with varied success. This work presents a representation-agnostic formulation of BRL under partially observability, unifying the previous models under one theoretical umbrella. To demonstrate its practical significance we also propose a novel derivation, Bayes-Adaptive Deep Dropout rl (BADDr), based on dropout networks. Under this parameterization, in contrast to previous work, the belief over the state and dynamics is a more scalable inference problem. We choose actions through Monte-Carlo tree search and empirically show that our method is competitive with state-of-the-art BRL methods on small domains while being able to solve much larger ones.