A distributional code for value in dopamine-based reinforcement learning

A distributional code for value in dopamine-based reinforcement learning
复制标题

DOI:
10.1038/s41586-019-1924-6
复制
发表时间:
2020-01-15
期刊:
影响因子:
64.8
通讯作者:
Botvinick, Matthew
Botvinick, Matthew
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Dabney, Will;Kurth-Nelson, Zeb;Botvinick, Matthew

文献摘要

被引文献

相似文献

自提出以来,多巴胺的奖赏预测误差理论已经解释了大量的经验现象,为理解大脑中奖赏和价值的表征提供了一个统一的框架(1-3)。根据现在的规范理论,奖励预测被表示为单个标量,它支持学习随机结果的期望或平均值。在这里,我们提出了一个基于多巴胺的强化学习的帐户,灵感来自最近的人工智能研究分布式强化学习(4-6)。我们假设,大脑并不是将未来可能的奖励表现为单一的平均值,而是表现为一种概率分布,有效地同时并行地表现出多个未来结果。这个想法意味着一组经验的预测,我们测试使用单单位记录从小鼠腹侧被盖区。我们的研究结果为分布式强化学习的神经实现提供了强有力的证据。
Since its introduction, the reward prediction error theory of dopamine has explained a wealth of empirical phenomena, providing a unifying framework for understanding the representation of reward and value in the brain(1-3). According to the now canonical theory, reward predictions are represented as a single scalar quantity, which supports learning about the expectation, or mean, of stochastic outcomes. Here we propose an account of dopamine-based reinforcement learning inspired by recent artificial intelligence research on distributional reinforcement learning(4-6). We hypothesized that the brain represents possible future rewards not as a single mean, but instead as a probability distribution, effectively representing multiple future outcomes simultaneously and in parallel. This idea implies a set of empirical predictions, which we tested using single-unit recordings from mouse ventral tegmental area. Our findings provide strong evidence for a neural realization of distributional reinforcement learning.