Temporal difference models and reward-related learning in the human brain

Temporal difference models and reward-related learning in the human brain
复制标题

DOI:
10.1016/s0896-6273(03)00169-7
复制
发表时间:
2003-04-24
期刊:
影响因子:
16.2
通讯作者:
Dolan, RJ
Dolan, RJ
中科院分区:
医学1区
文献类型:
--
作者:
O'Doherty, JP;Dayan, P;Dolan, RJ

文献摘要

被引文献

相似文献

时间差异学习被认为是巴甫洛夫条件反射的一个模型,在巴甫洛夫条件反射中,动物学会在条件刺激(CS)出现后预测奖励的传递。这个模型的一个关键组成部分是一个预测误差信号,在学习之前,它在奖励出现的时间做出反应,但在学习之后,它的反应会转移到CS开始的时间。为了测试显示这一信号的区域,研究人员使用与事件相关的功能磁共振成像对受试者进行扫描,同时给予愉悦的味觉奖励。回归分析显示,腹侧纹状体和眶额皮质的反应与该错误信号显著相关,表明在食欲条件反射过程中,由时间差异学习描述的计算在人脑中得到表达。
Temporal difference learning has been proposed as a model for Pavlovian conditioning, in which an animal learns to predict delivery of reward following presentation of a conditioned stimulus (CS). A key component of this model is a prediction error signal, which, before learning, responds at the time of presentation of reward but, after learning, shifts its response to the time of onset of the CS. In order to test for regions manifesting this signal profile, subjects were scanned using event-related fMRI while undergoing appetitive conditioning with a pleasant taste reward. Regression analyses revealed that responses in ventral striatum and orbitofrontal cortex were significantly correlated with this error signal, suggesting that, during appetitive conditioning, computations described by temporal difference learning are expressed in the human brain.