Dopamine cells respond to predicted events during classical conditioning: Evidence for eligibility traces in the reward-learning network

Dopamine cells respond to predicted events during classical conditioning: Evidence for eligibility traces in the reward-learning network
复制标题

DOI:
10.1523/jneurosci.1478-05.2005
复制
发表时间:
2005-06-29
影响因子:
5.3
通讯作者:
Hyland, BI
Hyland, BI
中科院分区:
医学1区
文献类型:
--
作者:
Pan, WX;Schmidt, R;Hyland, BI

文献摘要

被引文献

相似文献

提示-奖赏配对的行为条件反射导致中脑多巴胺(DA)细胞活动从对奖赏的反应转变为对预测性提示的反应。然而,这一转变背后的确切时间进程和机制仍不清楚。在这里,我们报告了一个组合的单单元记录和时间差(TD)建模方法来解决这个问题。来自清醒大鼠的记录数据表明,DA细胞在对条件提示的反应已经形成后,至少在训练的早期,保留了对预测奖励的反应。这与以前的TD模型,预测一个逐步逐步转移的潜伏期与奖励的反应失去了反应之前,发展到条件提示。通过探索TD参数空间,我们证明了在条件反射过程中DAcells的持续奖励反应只能通过具有长期资格痕迹(参数λ的非零值)和低学习率(alpha)的TD模型准确复制。TD参数的这些生理限制表明,资格痕迹和低的每次试验率的塑料修改可能是大脑中奖励学习的神经回路的基本特征。当刺激-奖励配对的数量有限时,这些属性使学习能够快速但稳定地启动,从而在现实世界环境中赋予显着的适应性优势。
Behavioral conditioning of cue-reward pairing results in a shift of midbrain dopamine (DA) cell activity from responding to the reward to responding to the predictive cue. However, the precise time course and mechanism underlying this shift remain unclear. Here, we report a combined single-unit recording and temporal difference (TD) modeling approach to this question. The data from recordings in conscious rats showed that DA cells retain responses to predicted reward after responses to conditioned cues have developed, at least early in training. This contrasts with previous TD models that predict a gradual stepwise shift in latency with responses to rewards lost before responses develop to the conditioned cue. By exploring the TD parameter space, we demonstrate that the persistent reward responses of DAcells during conditioning are only accurately replicated by a TD model with long-lasting eligibility traces (nonzero values for the parameter lambda) and low learning rate (alpha). These physiological constraints for TD parameters suggest that eligibility traces and low per-trial rates of plastic modification may be essential features of neural circuits for reward learning in the brain. Such properties enable rapid but stable initiation of learning when the number of stimulus-reward pairings is limited, conferring significant adaptive advantages in real-world environments.