Tonically active neurons in the striatum differentiate between delivery and omission of expected reward in a probabilistic task context

Tonically active neurons in the striatum differentiate between delivery and omission of expected reward in a probabilistic task context
复制标题

DOI:
10.1111/j.1460-9568.2009.06872.x
复制
发表时间:
2009-08-01
影响因子:
3.4
通讯作者:
Legallet, Eric
Legallet, Eric
中科院分区:
医学3区
文献类型:
--
作者:
Apicella, Paul;Deffains, Marc;Legallet, Eric

文献摘要

被引文献

相似文献

灵长类纹状体中的张力活跃神经元(TANs)对奖励刺激有反应,它们被认为与刺激-奖励关联或习惯的储存有关。然而,目前尚不清楚这些神经元是否可能在纹状体水平上作为奖励预测误差的可能神经元关联来发出奖励预测与其实际结果之间的差异的信号。为了解决这个问题,我们研究了在经典条件反射任务中训练的三只猴子的TANs的活动,在经典条件反射任务中,在视觉刺激之前进行液体奖励,奖励概率在不同的试验块之间系统地变化。在刺激-奖励间隔期间,通过监测猴子的嘴部运动来评估猴子根据概率区分条件的能力。我们发现,当奖励概率降低时,典型的TAN暂停反应显著增强,而当奖励概率高时,对预测刺激的反应略强。此外,TAN对奖励缺失的反应包括活动的减少或增加,随着奖励可能性的增加而变得更强。因此,似乎一组神经元对奖励传递和奖励遗漏的反应不同,活动方向相反,而另一组神经元的反应方向相同。这些数据表明,只有一部分TANs能够检测到奖励与预测不同的程度,从而有助于编码与强化学习相关的正奖励和负奖励预测误差。
Tonically active neurons (TANs) in the primate striatum are responsive to rewarding stimuli and they are thought to be involved in the storage of stimulus-reward associations or habits. However, it is unclear whether these neurons may signal the difference between the prediction of reward and its actual outcome as a possible neuronal correlate of reward prediction errors at the striatal level. To address this question, we studied the activity of TANs from three monkeys trained in a classical conditioning task in which a liquid reward was preceded by a visual stimulus and reward probability was systematically varied between blocks of trials. The monkeys' ability to discriminate the conditions according to probability was assessed by monitoring their mouth movements during the stimulus-reward interval. We found that the typical TAN pause responses to the delivery of reward were markedly enhanced as the probability of reward decreased, whereas responses to the predictive stimulus were somewhat stronger for high reward probability. In addition, TAN responses to the omission of reward consisted of either decreases or increases in activity that became stronger with increasing reward probability. It therefore appears that one group of neurons differentially responded to reward delivery and reward omission with changes in activity into opposite directions, while another group responded in the same direction. These data indicate that only a subset of TANs could detect the extent to which reward occurs differently than predicted, thus contributing to the encoding of positive and negative reward prediction errors that is relevant to reinforcement learning.