Dopamine neurons learn to encode the long-term value of multiple future rewards

Dopamine neurons learn to encode the long-term value of multiple future rewards
复制标题

DOI:
10.1073/pnas.1014457108
复制
发表时间:
2011-09-13
影响因子:
11.1
通讯作者:
Kimura, Minoru
Kimura, Minoru
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Enomoto, Kazuki;Matsumoto, Naoyuki;Kimura, Minoru

文献摘要

被引文献

相似文献

中脑多巴胺神经元信号奖励价值,他们的预测误差,和事件的显着性。如果它们在实现特定的远期目标方面发挥了关键作用,那么未来的长期奖励也应该像强化学习理论所建议的那样被编码。在这里,我们解决这个未经实验检验的问题。我们记录了三只猴子的185个多巴胺神经元,这些猴子执行多步选择任务,在这些任务中,它们在备选方案中探索一个奖励目标,然后利用这些知识在随后的一系列试验中选择相同的目标来获得一个或两个额外的奖励。一项对舔水行为的分析表明,猴子在单独的试验中并没有预料到立即预期的奖励,而是预料到了立即和多次未来奖励的总和。根据这种行为观察,多巴胺对开始提示和提示器哔哔声的反应分别反映了多个未来奖励及其错误的预期值。更具体地说,当猴子在几周内学习多步选择任务时,多巴胺神经元的反应编码了立即和预期的多种未来奖励的总和。通过对强化学习中时间折扣价值函数的理论描述,定量预测多巴胺反应。这些发现表明,多巴胺神经元学会对多种未来奖励的长期价值进行编码,而对遥远的奖励进行打折。
Midbrain dopamine neurons signal reward value, their prediction error, and the salience of events. If they play a critical role in achieving specific distant goals, long-term future rewards should also be encoded as suggested in reinforcement learning theories. Here, we address this experimentally untested issue. We recorded 185 dopamine neurons in three monkeys that performed a multistep choice task in which they explored a reward target among alternatives and then exploited that knowledge to receive one or two additional rewards by choosing the same target in a set of subsequent trials. An analysis of anticipatory licking for reward water indicated that the monkeys did not anticipate an immediately expected reward in individual trials; rather, they anticipated the sum of immediate and multiple future rewards. In accordance with this behavioral observation, the dopamine responses to the start cues and reinforcer beeps reflected the expected values of the multiple future rewards and their errors, respectively. More specifically, when monkeys learned the multistep choice task over the course of several weeks, the responses of dopamine neurons encoded the sum of the immediate and expected multiple future rewards. The dopamine responses were quantitatively predicted by theoretical descriptions of the value function with time discounting in reinforcement learning. These findings demonstrate that dopamine neurons learn to encode the long-term value of multiple future rewards with distant rewards discounted.