Emergence of belief-like representations through reinforcement learning.

Emergence of belief-like representations through reinforcement learning.
复制标题

通过强化学习出现类似信念的表征。

DOI:
10.1101/2023.04.04.535512
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Gershman,SamuelJ
Gershman,SamuelJ
中科院分区:
--
文献类型:
--
作者:
Hennig,JayA;Pinto,SandraARomero;Yamaguchi,Takahiro;Linderman,ScottW;Uchida,Naoshige;Gershman,SamuelJ

文献摘要

相似文献

为了适应,动物必须学会预测未来的回报或价值。为了做到这一点,动物被认为使用强化学习来学习奖励预测。然而,与经典模型相反,动物必须学会仅使用不完整的状态信息来估计价值。以前的工作表明,动物估计值部分可观察的任务,首先形成“信念”-最佳贝叶斯估计的隐藏状态的任务。虽然这是解决部分可观测性问题的一种方法,但它不是唯一的方法,也不是复杂现实环境中计算可扩展性最强的解决方案。在这里,我们证明了递归神经网络(RNN)可以直接从观察中学习估计值,产生类似于实验观察到的奖励预测错误,而没有任何明确的估计信念的目标。我们整合了统计、功能和动力系统对信念的观点,以表明RNN的学习表示对信念信息进行编码,但只有当RNN的容量足够大时。这些结果说明了动物如何在没有明确估计信念的情况下估计任务中的价值,从而产生对容量有限的系统有用的表示。
To behave adaptively, animals must learn to predict future reward, or value. To do this, animals are thought to learn reward predictions using reinforcement learning. However, in contrast to classical models, animals must learn to estimate value using only incomplete state information. Previous work suggests that animals estimate value in partially observable tasks by first forming “beliefs”—optimal Bayesian estimates of the hidden states in the task. Although this is one way to solve the problem of partial observability, it is not the only way, nor is it the most computationally scalable solution in complex, real-world environments. Here we show that a recurrent neural network (RNN) can learn to estimate value directly from observations, generating reward prediction errors that resemble those observed experimentally, without any explicit objective of estimating beliefs. We integrate statistical, functional, and dynamical systems perspectives on beliefs to show that the RNN’s learned representation encodes belief information, but only when the RNN’s capacity is sufficiently large. These results illustrate how animals can estimate value in tasks without explicitly estimating beliefs, yielding a representation useful for systems with limited capacity.