Modelling individual differences in the form of Pavlovian conditioned approach responses: a dual learning systems approach with factored representations.

Modelling individual differences in the form of Pavlovian conditioned approach responses: a dual learning systems approach with factored representations.
复制标题

DOI:
10.1371/journal.pcbi.1003466
复制
发表时间:
2014-02
影响因子:
4.3
通讯作者:
Khamassi M
Khamassi M
中科院分区:
生物学2区
文献类型:
--
作者:
Lesaint F;Sigaud O;Flagel SB;Robinson TE;Khamassi M

文献摘要

参考文献

被引文献

相似文献

强化学习极大地影响了条件反射的模型,为后天习得行为和潜在的生理观察提供了强有力的解释。然而,在最近的大鼠自整形实验中,巴甫洛夫条件反应(CRS)形式的变化及其相关的多巴胺活性对经典假说提出了质疑,即相性多巴胺活性对应于来自经典无模型系统的奖赏预测误差样信号,这是巴甫洛夫条件反射所必需的。在使用食物作为无条件刺激的巴甫洛夫条件反射过程中,一些大鼠(符号跟踪者)越来越热衷于接近和接触条件刺激(CS)本身--一个杠杆--而另一些大鼠(目标跟踪者)则在条件刺激(CS)呈现时学习接近食物传递的位置。重要的是,尽管信号跟踪器和目标跟踪器都很好地了解了CS-US的联系,但只有在信号跟踪器中,阶段多巴胺活动才会显示出典型的奖赏预测误差样的爆发。此外,目标跟踪CR的获得和表达都不依赖于多巴胺。在这里,我们提出了一个计算模型,可以解释这种个体差异。我们表明,基于模型的系统和修订的无模型系统的组合可以解释大鼠不同的CRS的发展。此外,我们表明,修改一个经典的无模型系统,通过使用因式分解表示法来单独处理刺激,可以解释为什么在一些大鼠身上可以观察到经典的多巴胺能模式,而在另一些大鼠身上则不能,这取决于它们发展的CR。此外,该模型可以解释使用相同或类似的自动整形过程获得的其他行为和药理学结果。最后,该模型使得绘制一组实验预测成为可能,这些实验预测可以在修改后的实验方案中得到验证。我们建议,在计算神经科学研究中进一步研究因式表示可能是有用的。对奖赏的完全预测因素的反应,即巴甫洛夫条件反射,长期以来一直被用强化学习理论来解释。这一理论使学习过程正式化,通过将价值归因于情景和行动,使其有可能将行为引向奖励目标。有趣的是,在这类实验中,隐含的机制依赖于与多巴胺神经元活动平行的强化信号。然而,最近的研究挑战了用单一过程解释巴甫洛夫条件反射的经典观点。当拿到一个在运送食物之前被收回的杠杆时,一些老鼠开始咀嚼和咬食物杂志,而另一些老鼠则开始咀嚼和咬杠杆,即使没有必要相互作用来获得食物。这些差异在大脑活动中也很明显,当用药物测试时,表明多种系统共存。我们提出了一个计算模型,扩展了经典理论来解释这些数据。有趣的是,我们可以从这个模型中得出可能得到实验验证的预测。受用于模拟工具性行为(需要采取行动才能获得回报)和高级巴甫洛夫行为(如过度期望、消极模式)的机制的启发,它为开始模拟观察到的它们之间的强烈互动提供了一个切入点。
Reinforcement Learning has greatly influenced models of conditioning, providing powerful explanations of acquired behaviour and underlying physiological observations. However, in recent autoshaping experiments in rats, variation in the form of Pavlovian conditioned responses (CRs) and associated dopamine activity, have questioned the classical hypothesis that phasic dopamine activity corresponds to a reward prediction error-like signal arising from a classical Model-Free system, necessary for Pavlovian conditioning. Over the course of Pavlovian conditioning using food as the unconditioned stimulus (US), some rats (sign-trackers) come to approach and engage the conditioned stimulus (CS) itself – a lever – more and more avidly, whereas other rats (goal-trackers) learn to approach the location of food delivery upon CS presentation. Importantly, although both sign-trackers and goal-trackers learn the CS-US association equally well, only in sign-trackers does phasic dopamine activity show classical reward prediction error-like bursts. Furthermore, neither the acquisition nor the expression of a goal-tracking CR is dopamine-dependent. Here we present a computational model that can account for such individual variations. We show that a combination of a Model-Based system and a revised Model-Free system can account for the development of distinct CRs in rats. Moreover, we show that revising a classical Model-Free system to individually process stimuli by using factored representations can explain why classical dopaminergic patterns may be observed for some rats and not for others depending on the CR they develop. In addition, the model can account for other behavioural and pharmacological results obtained using the same, or similar, autoshaping procedures. Finally, the model makes it possible to draw a set of experimental predictions that may be verified in a modified experimental protocol. We suggest that further investigation of factored representations in computational neuroscience studies may be useful. Acquisition of responses towards full predictors of rewards, namely Pavlovian conditioning, has long been explained using the reinforcement learning theory. This theory formalizes learning processes that, by attributing values to situations and actions, makes it possible to direct behaviours towards rewarding objectives. Interestingly, the implied mechanisms rely on a reinforcement signal that parallels the activity of dopamine neurons in such experiments. However, recent studies challenged the classical view of explaining Pavlovian conditioning with a single process. When presented with a lever whose retraction preceded the delivery of food, some rats started to chew and bite the food magazine whereas others chew and bite the lever, even if no interactions were necessary to get the food. These differences were also visible in brain activity and when tested with drugs, suggesting the coexistence of multiple systems. We present a computational model that extends the classical theory to account for these data. Interestingly, we can draw predictions from this model that may be experimentally verified. Inspired by mechanisms used to model instrumental behaviours, where actions are required to get rewards, and advanced Pavlovian behaviours (such as overexpectation, negative patterning), it offers an entry point to start modelling the strong interactions observed between them.
DOI: 10.1016/j.bbr.2012.02.032
发表时间: 2012-05-01
影响因子: 2.7
作者:
DiFeliceantonio, Alexandra G.;Berridge, Kent C.
通讯作者: Berridge, Kent C.
DOI: 10.1016/j.neuron.2011.02.027
发表时间: 2011-03-24
期刊: Neuron
影响因子: 16.2
作者:
Daw ND;Gershman SJ;Seymour B;Dayan P;Dolan RJ
通讯作者: Dolan RJ
DOI: 10.3389/fnbot.2013.00003
发表时间: 2013
影响因子: 3.1
作者:
Elfwing S;Uchibe E;Doya K
通讯作者: Doya K
DOI: 10.1023/a:1008965713435
发表时间: 1999-07-01
期刊: AUTONOMOUS ROBOTS
影响因子: 3.5
作者:
Balkenius, C;Morén, J
通讯作者: Morén, J
DOI: 10.3758/bf03209705
发表时间: 1979-01-01
期刊: ANIMAL LEARNING & BEHAVIOR
影响因子: --
作者:
BALSAM, PD;PAYNE, D
通讯作者: PAYNE, D