Reinforcement Learning and Episodic Memory in Humans and Animals: An Integrative Framework.

Reinforcement Learning and Episodic Memory in Humans and Animals: An Integrative Framework.
复制标题

DOI:
10.1146/annurev-psych-122414-033625
复制
发表时间:
2017-01-03
影响因子:
24.8
通讯作者:
Daw ND
Daw ND
中科院分区:
心理学1区
文献类型:
--
作者:
Gershman SJ;Daw ND

文献摘要

参考文献

被引文献

相似文献

我们回顾了强化学习(RL)的心理学和神经科学,在过去的二十年里,通过对简单学习和决策任务的全面实验研究,强化学习取得了重大进展。然而,这些任务的简单性忽略了强化学习在真实的世界中的重要方面:(i)状态空间是高维的,连续的,部分可观察的;这意味着(ii)数据相对稀疏:实际上完全相同的情况可能永远不会遇到两次;以及(iii)奖励取决于行为的长期后果,其方式违反了使RL易于处理的经典假设。一个看似独特的挑战是,从认知角度来看,这些理论在很大程度上与程序记忆和语义记忆有关:从许多经验中逐渐提取的关于行动价值或世界模型的知识如何驱动选择。这忽略了与个体事件痕迹相关的记忆的许多方面,例如情景记忆。我们认为,这两个差距是相关的。特别是,计算挑战可以通过赋予RL系统情景记忆来部分解决,使它们能够(i)在复杂的状态空间上有效地近似值函数,(ii)用很少的数据学习,以及(iii)在行动和奖励之间建立长期依赖关系。我们回顾了支持这一建议的计算理论和经验证据,我们的建议表明,在RL中无处不在的和不同的角色记忆可能作为一个综合学习系统的一部分。
We review the psychology and neuroscience of reinforcement learning (RL), which has witnessed significant progress in the last two decades, enabled by the comprehensive experimental study of simple learning and decision-making tasks. However, the simplicity of these tasks misses important aspects of reinforcement learning in the real world: (i) State spaces are high-dimensional, continuous, and partially observable; this implies that (ii) data are relatively sparse: indeed precisely the same situation may never be encountered twice; and also that (iii) rewards depend on long-term consequences of actions in ways that violate the classical assumptions that make RL tractable. A seemingly distinct challenge is that, cognitively, these theories have largely connected with procedural and semantic memory: how knowledge about action values or world models extracted gradually from many experiences can drive choice. This misses many aspects of memory related to traces of individual events, such as episodic memory. We suggest that these two gaps are related. In particular, the computational challenges can be dealt with, in part, by endowing RL systems with episodic memory, allowing them to (i) efficiently approximate value functions over complex state spaces, (ii) learn with very little data, and (iii) bridge long-term dependencies between actions and rewards. We review the computational theory underlying this proposal and the empirical evidence to support it. Our proposal suggests that the ubiquitous and diverse roles of memory in RL may function as part of an integrated learning system.
神经元型特异性信号,用于腹侧对段区域的奖励和惩罚。
DOI: 10.1038/nature10754
发表时间: 2012-01-18
期刊: NATURE
影响因子: 64.8
作者:
Cohen, Jeremiah Y.;Haesler, Sebastian;Vong, Linh;Lowell, Bradford B.;Uchida, Naoshige
通讯作者: Uchida, Naoshige
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.1016/j.jmp.2008.05.006
发表时间: 2009-06-01
影响因子: 1.8
作者:
Biele, Guido;Erev, Ido;Ert, Eyal
通讯作者: Ert, Eyal
海马和纹状体系统之间的合作相互作用支持灵活导航。
DOI: 10.1016/j.neuroimage.2012.01.046
发表时间: 2012-04-02
期刊: NEUROIMAGE
影响因子: 5.7
作者:
Brown, Thackery I.;Ross, Robert S.;Tobyne, Sean M.;Stern, Chantal E.
通讯作者: Stern, Chantal E.
DOI: 10.1016/j.cognition.2008.08.011
发表时间: 2009-12
期刊: Cognition
影响因子: 3.4
作者:
Botvinick MM;Niv Y;Barto AG
通讯作者: Barto AG