Dopamine-dependent prediction errors underpin reward-seeking behaviour in humans.

Dopamine-dependent prediction errors underpin reward-seeking behaviour in humans.
复制标题

DOI:
10.1038/nature05051
复制
发表时间:
2006-08-31
期刊:
影响因子:
64.8
通讯作者:
Frith, Chris D
Frith, Chris D
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Pessiglione, Mathias;Seymour, Ben;Flandin, Guillaume;Dolan, Raymond J;Frith, Chris D

文献摘要

被引文献

相似文献

工具性学习理论的核心是理解如何利用成功和失败来改进未来的决策。这些理论强调了奖励预测错误在更新与可用操作相关的值方面的核心作用。在动物中,大量证据表明,神经递质多巴胺通过调节皮质纹状体突触功效的能力,可能在此类学习中发挥关键作用。然而,没有直接证据将多巴胺、纹状体活动和人类的行为选择联系起来。在这里,我们表明,在工具学习过程中,纹状体中表达的奖励预测误差的大小是通过服用增强(3,4-二羟基-l-苯丙氨酸;l-DOPA)或减少(氟哌啶醇)多巴胺能功能的药物来调节的。因此,相对于用氟哌啶醇治疗的受试者,用l-DOPA治疗的受试者更倾向于选择最有益的行动。此外,将预测误差的大小纳入标准的行动价值学习算法可以准确地再现受试者在不同药物条件下的行为选择。我们得出的结论是,纹状体活动的多巴胺依赖性调节可以解释人脑如何利用奖励预测错误来改善未来的决策。
Theories of instrumental learning are centred on understanding how success and failure are used to improve future decisions. These theories highlight a central role for reward prediction errors in updating the values associated with available actions. In animals, substantial evidence indicates that the neurotransmitter dopamine might have a key function in this type of learning, through its ability to modulate cortico-striatal synaptic efficacy. However, no direct evidence links dopamine, striatal activity and behavioural choice in humans. Here we show that, during instrumental learning, the magnitude of reward prediction error expressed in the striatum is modulated by the administration of drugs enhancing (3,4-dihydroxy-l-phenylalanine;l-DOPA) or reducing (haloperidol) dopaminergic function. Accordingly, subjects treated withl-DOPA have a greater propensity to choose the most rewarding action relative to subjects treated with haloperidol. Furthermore, incorporating the magnitude of the prediction errors into a standard action-value learning algorithm accurately reproduced subjects' behavioural choices under the different drug conditions. We conclude that dopamine-dependent modulation of striatal activity can account for how the human brain uses reward prediction errors to improve future decisions.