Dopamine Ramps Are a Consequence of Reward Prediction Errors

Dopamine Ramps Are a Consequence of Reward Prediction Errors
复制标题

DOI:
10.1162/neco_a_00559
复制
发表时间:
2014-03-01
期刊:
影响因子:
2.9
通讯作者:
Gershman, Samuel J.
Gershman, Samuel J.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Gershman, Samuel J.

文献摘要

被引文献

相似文献

多巴胺的时间差异学习模型断言,多巴胺的阶段水平编码奖励预测误差。然而,这一假设受到了最近观察的挑战,随着目标的接近,层多巴胺水平逐渐上升。本文描述了时间差异学习模型预测多巴胺激增的条件。关键思想是代表性的:接近目标的二次变换意味着近似线性斜坡,正如实验观察到的那样。
Temporal difference learning models of dopamine assert that phasic levels of dopamine encode a reward prediction error. However, this hypothesis has been challenged by recent observations of gradually ramping stratal dopamine levels as a goal is approached. This note describes conditions under which temporal difference learning models predict dopamine ramping. The key idea is representational: a quadratic transformation of proximity to the goal implies approximately linear ramping, as observed experimentally.