States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning.

States versus rewards: dissociable neural prediction error signals underlying model-based and model-free reinforcement learning.
复制标题

DOI:
10.1016/j.neuron.2010.04.016
复制
发表时间:
2010-05-27
期刊:
影响因子:
16.2
通讯作者:
O'Doherty, John P.
O'Doherty, John P.
中科院分区:
医学1区
文献类型:
--
作者:
Glaescher, Jan;Daw, Nathaniel;Dayan, Peter;O'Doherty, John P.

文献摘要

参考文献

被引文献

相似文献

强化学习(RL)利用对情境(“状态”)和结果的连续经验来评估行动。无模型RL以奖励预测误差(RPE)的形式直接使用这种经验,而基于模型的RL间接使用这种经验,建立环境的状态转换和结果结构的模型,并通过搜索该模型来评估行为。状态预测误差(SPE)起着核心作用,它报告当前模型和观察到的状态转换之间的差异。在人类解决概率马尔可夫决策任务时,使用功能磁共振成像,我们发现了顶内沟和外侧前额叶皮质的SPE的神经特征,以及以前很好描述的腹侧纹状体的RPE。这一发现支持人类存在两种独特的学习信号,这可能构成指导行为的不同计算策略的基础。
Reinforcement learning (RL) uses sequential experience with situations (“states”) and outcomes to assess actions. Whereas model-free RL uses this experience directly, in the form of a reward prediction error (RPE), model-based RL uses it indirectly, building a model of the state transition and outcome structure of the environment, and evaluating actions by searching this model. A state prediction error (SPE) plays a central role, reporting discrepancies between the current model and the observed state transitions. Using functional magnetic resonance imaging in humans solving a probabilistic Markov decision task we found the neural signature of an SPE in the intraparietal sulcus and lateral prefrontal cortex, in addition to the previously well-characterized RPE in the ventral striatum. This finding supports the existence of two unique forms of learning signal in humans, which may form the basis of distinct computational strategies for guiding behavior.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1080/02724990042000010
发表时间: 2001-02-01
期刊: QUARTERLY JOURNAL OF EXPERIMENTAL PSYCHOLOGY SECTION B-COMPARATIVE AND PHYSIOLOGICAL PSYCHOLOGY
影响因子: --
作者:
Dickinson, A
通讯作者: Dickinson, A
DOI: 10.1162/089976602753712972
发表时间: 2002-06-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Doya, K;Samejima, K;Kawato, M
通讯作者: Kawato, M
DOI: 10.1038/73009
发表时间: 2000-03-01
影响因子: 25
作者:
Corbetta, M;Kincade, JM;Shulman, GL
通讯作者: Shulman, GL
DOI: 10.1007/s12021-008-9042-x
发表时间: 2009-03-01
期刊: NEUROINFORMATICS
影响因子: 3
作者:
Glaescher, Jan
通讯作者: Glaescher, Jan