Tamping Ramping: Algorithmic, Implementational, and Computational Explanations of Phasic Dopamine Signals in the Accumbens.

Tamping Ramping: Algorithmic, Implementational, and Computational Explanations of Phasic Dopamine Signals in the Accumbens.
复制标题

DOI:
10.1371/journal.pcbi.1004622
复制
发表时间:
2015-12
影响因子:
4.3
通讯作者:
Dayan P
Dayan P
中科院分区:
生物学2区
文献类型:
--
作者:
Lloyd K;Dayan P

文献摘要

被引文献

相似文献

大量证据表明,多巴胺神经元的相活动代表了强化学习的时间差异预测误差。然而,最近有报道称,当动物即将采取行动或即将获得奖励时,纹状体中的多巴胺浓度会陡坡式增加,这似乎对既定思维构成了挑战。这是因为隐含的活动可以被先前的刺激持续预测,因此不会出现这种预测误差。在这里,我们探讨了这种斜坡信号的三种可能的解释:(a)解决行动时间的不确定性;(b)多巴胺对与做出选择有关的机制的直接影响;(三)活力打折的新模式。总的来说,这些都表明,多巴胺的增加可以用标准的理论来解释,只有轻微的干扰,尽管关于其近端原因仍然存在迫切的问题。我们建议用实验方法来解开哪一种机制是导致多巴胺增加的原因。长期以来,多巴胺一直与奖励动机行为有关。理论和实验表明,含有多巴胺的神经元的活动类似于一种用于学习对未来奖励的预期的时间复杂的预测错误。这种解释似乎与最近对“斜坡”的观察不一致,即在执行行动或获得奖励之前,细胞外多巴胺浓度逐渐增加。我们探讨了三种不同的可能解释这种上升信号的产生:(a)当受试者对何时执行行动感到不确定时;(b)多巴胺本身影响选择的时间进程;(c)在一种新的模式下,“准强直性”多巴胺信号通过一种形式的时间折扣产生。因此,我们表明多巴胺斜坡可以与当前的理论相结合,并建议通过实验来阐明所涉及的机制。
Substantial evidence suggests that the phasic activity of dopamine neurons represents reinforcement learning’s temporal difference prediction error. However, recent reports of ramp-like increases in dopamine concentration in the striatum when animals are about to act, or are about to reach rewards, appear to pose a challenge to established thinking. This is because the implied activity is persistently predictable by preceding stimuli, and so cannot arise as this sort of prediction error. Here, we explore three possible accounts of such ramping signals: (a) the resolution of uncertainty about the timing of action; (b) the direct influence of dopamine over mechanisms associated with making choices; and (c) a new model of discounted vigour. Collectively, these suggest that dopamine ramps may be explained, with only minor disturbance, by standard theoretical ideas, though urgent questions remain regarding their proximal cause. We suggest experimental approaches to disentangling which of the proposed mechanisms are responsible for dopamine ramps. Dopamine has long been implicated in reward-motivated behaviour. Theory and experiments suggest that activity of dopamine-containing neurons resembles a temporally-sophisticated prediction error used to learn expectations of future reward. This account would appear to be inconsistent with recent observations of ‘ramps’, i.e., gradual increases in extracellular dopamine concentration prior to the execution of actions or the acquisition of rewards. We explore three different possible explanations of such ramping signals as arising: (a) when subjects experience uncertainty about when actions will be executed; (b) when dopamine itself influences the timecourse of choice; and (c) under a new model in which ‘quasi-tonic’ dopamine signals arise through a form of temporal discounting. We thereby show that dopamine ramps can be integrated with current theories, and also suggest experiments to clarify which mechanisms are involved.