Forgetting in Reinforcement Learning Links Sustained Dopamine Signals to Motivation.

Forgetting in Reinforcement Learning Links Sustained Dopamine Signals to Motivation.
复制标题

DOI:
10.1371/journal.pcbi.1005145
复制
发表时间:
2016-10
影响因子:
4.3
通讯作者:
Morita K
Morita K
中科院分区:
生物学2区
文献类型:
--
作者:
Kato A;Morita K

文献摘要

参考文献

被引文献

相似文献

有人认为多巴胺(DA)代表强化学习中定义的奖励-预测-错误(RPE),因此多巴胺对不可预测的奖励做出反应,而不是预测的奖励。然而,最近的研究发现,在涉及自我节奏行为的任务中,DA反应对可预测的奖励持续存在,并表明这种反应代表了一种动机信号。我们之前已经证明,如果存在学习值的衰减/遗忘,RPE可以维持,这可以通过存储学习值的突触强度的衰减来实现。然而,这一解释并没有解释强直/持续DA和动机之间的联系。在目前的工作中,我们探索了价值衰减在自定节奏接近行为中的动机效应,建模为一系列“Go”或“No-Go”的目标选择。通过模拟,我们发现价值衰减可以增强动机,特别是促进快速达成目标,尽管这与直觉相反。数学分析表明,潜在的潜在机制有两方面:(1)衰减诱导的持续RPE创造了朝向目标的“Go”值梯度;(2)“Go”和“No-Go”之间的价值对比产生了,因为当选择的值不断更新时,未选择的值只是衰减。我们的模型为表明DA在动机中的作用的关键实验结果提供了潜在的解释:(i)通过训练后阻断DA信号导致的行为放缓,(ii)观察到DA阻断严重损害了获得奖励的努力行为,同时在很大程度上避免了对容易获得的奖励的追求,以及(iii)奖励数量、行为速度所反映的动机水平和平均DA水平之间的关系。这些结果表明,具有价值衰减或遗忘的强化学习为数据处理在价值学习和动机中的作用提供了一个简洁的机制解释。我们的研究结果还表明,即使学习明显趋同,当价值学习的生物系统仍处于活跃状态时,系统可能处于动态平衡状态,即学习和遗忘处于平衡状态。多巴胺(DA)被认为有两个与奖励相关的角色:(1)表示奖励预测错误(RPE),(2)提供动机驱动。角色(1)基于生理结果,即DA对不可预测而非可预测的奖励做出反应,而角色(2)则得到药理学结果的支持,即DA信号的阻断会导致动机障碍,如自定节奏行为的减慢。到目前为止,这两个角色被认为是由两种不同的时间模式的DA信号发挥作用:作用(1)由相位信号和作用(2)由强音/持续信号。然而,最近的研究发现,持续的DA信号具有表明这两种作用(1)和(2)的特征,这使情况变得复杂。与此同时,虽然作用(1)的突触/回路机制,即RPE如何在DA神经元的上游计算,以及RPE依赖的学习值更新如何通过DA依赖的突触可塑性发生,现在已经明确,但作用(2)的机制仍然不清楚。在这项工作中,我们在强化学习框架中通过一系列“Go”或“No-Go”选择来模拟自定节奏的行为,假设DA的作用(1),并证明了学习值的衰减/遗忘的结合,这可能是作为存储学习值的突触强度的衰减来实现的,为DA的两个角色及其各种时间模式提供了一个潜在的统一机制解释。
It has been suggested that dopamine (DA) represents reward-prediction-error (RPE) defined in reinforcement learning and therefore DA responds to unpredicted but not predicted reward. However, recent studies have found DA response sustained towards predictable reward in tasks involving self-paced behavior, and suggested that this response represents a motivational signal. We have previously shown that RPE can sustain if there is decay/forgetting of learned-values, which can be implemented as decay of synaptic strengths storing learned-values. This account, however, did not explain the suggested link between tonic/sustained DA and motivation. In the present work, we explored the motivational effects of the value-decay in self-paced approach behavior, modeled as a series of ‘Go’ or ‘No-Go’ selections towards a goal. Through simulations, we found that the value-decay can enhance motivation, specifically, facilitate fast goal-reaching, albeit counterintuitively. Mathematical analyses revealed that underlying potential mechanisms are twofold: (1) decay-induced sustained RPE creates a gradient of ‘Go’ values towards a goal, and (2) value-contrasts between ‘Go’ and ‘No-Go’ are generated because while chosen values are continually updated, unchosen values simply decay. Our model provides potential explanations for the key experimental findings that suggest DA's roles in motivation: (i) slowdown of behavior by post-training blockade of DA signaling, (ii) observations that DA blockade severely impairs effortful actions to obtain rewards while largely sparing seeking of easily obtainable rewards, and (iii) relationships between the reward amount, the level of motivation reflected in the speed of behavior, and the average level of DA. These results indicate that reinforcement learning with value-decay, or forgetting, provides a parsimonious mechanistic account for the DA's roles in value-learning and motivation. Our results also suggest that when biological systems for value-learning are active even though learning has apparently converged, the systems might be in a state of dynamic equilibrium, where learning and forgetting are balanced. Dopamine (DA) has been suggested to have two reward-related roles: (1) representing reward-prediction-error (RPE), and (2) providing motivational drive. Role(1) is based on the physiological results that DA responds to unpredicted but not predicted reward, whereas role(2) is supported by the pharmacological results that blockade of DA signaling causes motivational impairments such as slowdown of self-paced behavior. So far, these two roles are considered to be played by two different temporal patterns of DA signals: role(1) by phasic signals and role(2) by tonic/sustained signals. However, recent studies have found sustained DA signals with features indicative of both roles (1) and (2), complicating this picture. Meanwhile, whereas synaptic/circuit mechanisms for role(1), i.e., how RPE is calculated in the upstream of DA neurons and how RPE-dependent update of learned-values occurs through DA-dependent synaptic plasticity, have now become clarified, mechanisms for role(2) remain unclear. In this work, we modeled self-paced behavior by a series of ‘Go’ or ‘No-Go’ selections in the framework of reinforcement-learning assuming DA's role(1), and demonstrated that incorporation of decay/forgetting of learned-values, which is presumably implemented as decay of synaptic strengths storing learned-values, provides a potential unified mechanistic account for the DA's two roles, together with its various temporal patterns.
DOI: 10.1523/jneurosci.3490-14.2015
发表时间: 2015-01-21
影响因子: 5.3
作者:
Damodaran, Sriraman;Cressman, John R.;Blackwell, Kim T.
通讯作者: Blackwell, Kim T.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1037/a0037015
发表时间: 2014-07-01
影响因子: 5.4
作者:
Collins, Anne G. E.;Frank, Michael J.
通讯作者: Frank, Michael J.
DOI: 10.3389/fpg.2015.00229
发表时间: 2015-03-12
影响因子: 3.8
作者:
Dai, Junyi;Kerestes, Rebecca;Stout, Julie C.
通讯作者: Stout, Julie C.
DOI: 10.1371/journal.pcbi.1003640
发表时间: 2014-06
影响因子: 4.3
作者:
Brea J;Urbanczik R;Senn W
通讯作者: Senn W