Deep active inference as variational policy gradients

Deep active inference as variational policy gradients
复制标题

DOI:
10.1016/j.jmp.2020.102348
复制
发表时间:
2020-06-01
影响因子:
1.8
通讯作者:
Millidge, Beren
Millidge, Beren
中科院分区:
心理学4区
文献类型:
--
作者:
Millidge, Beren

文献摘要

被引文献

相似文献

主动推理是一种源自理论神经科学的理论,它将行动和规划视为贝叶斯推理问题,通过最小化单个量(变分自由能)来解决。该理论承诺对行动和感知进行统一的解释,并结合生物学上合理的过程理论。然而,尽管有这些潜在的优势,当前主动推理的实现只能处理较小的策略和状态空间,并且通常需要了解环境动态。在本文中,我们提出了一种新颖的深度主动推理算法,该算法使用深度神经网络作为灵活的函数逼近器来逼近关键密度,这使得我们的方法能够扩展到比文献中之前尝试的任何更大、更复杂的任务。我们在一系列 OpenAIGym 基准任务上展示了我们的方法,并获得了与常见强化学习基准相当的性能。此外,我们的算法与最大熵强化学习和策略梯度算法有相似之处,这揭示了主动推理框架和强化学习之间有趣的联系。 (C) 2020 Elsevier Inc. 保留所有权利。
Active Inference is a theory arising from theoretical neuroscience which casts action and planning as Bayesian inference problems to be solved by minimizing a single quantity - the variational free energy. The theory promises a unifying account of action and perception coupled with a biologically plausible process theory. However, despite these potential advantages, current implementations of Active Inference can only handle small policy and state-spaces and typically require the environmental dynamics to be known. In this paper we propose a novel deep Active Inference algorithm that approximates key densities using deep neural networks as flexible function approximators, which enables our approach to scale to significantly larger and more complex tasks than any before attempted in the literature. We demonstrate our method on a suite of OpenAIGym benchmark tasks and obtain performance comparable with common reinforcement learning baselines. Moreover, our algorithm evokes similarities with maximum-entropy reinforcement learning and the policy gradients algorithm, which reveals interesting connections between the Active Inference framework and reinforcement learning. (C) 2020 Elsevier Inc. All rights reserved.