Learning from Delayed Reward und Punishment in a Spiking Neural Network Model of Basal Ganglia with Opposing D1/D2 Plasticity

Learning from Delayed Reward und Punishment in a Spiking Neural Network Model of Basal Ganglia with Opposing D1/D2 Plasticity
复制标题

从具有相反 D1/D2 可塑性的基底神经节尖峰神经网络模型中的延迟奖励和惩罚中学习

DOI:
10.1007/978-3-642-33269-2_58
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Tittgemeyer
Tittgemeyer
中科院分区:
--
文献类型:
--
作者:
Jitsev;Abraham;Morrison;Tittgemeyer

文献摘要

参考文献

相似文献

在之前工作的基础上,我们引入了基底节区从奖惩中学习的尖峰行为者-批评者网络模型。在该模型中,纹状体被分离成携带D1或D2多巴胺受体类型的中棘神经元(msn)群体。这种隔离允许在各自的人群中明确表示积极和消极的预期结果。根据最近的实验,我们进一步假设D1和D2 MSN群体具有相反的多巴胺调节的双向突触可塑性。实验是在一个网格世界中进行的,在这个世界中,一个移动的代理必须达到一个遥远的奖励目标状态。与之前的模型相反,网络不仅学会了接近奖励目标,还因此避免了惩罚。刺突网络模型解释了纹状体内D1/D2微MSN分离的功能作用,特别是在不同微MSN聚集的突触上发现的多巴胺依赖可塑性的反向。
Extending previous work, we introduce a spiking actor-critic network model of learning from reward and punishment in the basal ganglia. In the model, the striatum is taken to be segregated into populations of medium spiny neurons (MSNs) that carry either D1 or D2 dopamine receptor type. This segregation allows explicit representation of both positive and negative expected outcome within the respective population. In line with recent experiments, we further assume that D1 and D2 MSN populations have opposing dopamine-modulated bidirectional synaptic plasticity. Experiments were conducted in a grid world, where a moving agent had to reach a remote rewarded goal state. The network learned not only to approach the rewarded goal, but also to consequently avoid punishments as opposed to the previous model. The spiking network model explains functional role of D1/D2 MSN segregation within striatum, specifically the reversed direction of dopamine-dependent plasticity found at synapses converging on different MSNs.
DOI: 10.2976/1.2732246/10.2976/1
发表时间: 2007-05-01
期刊: HFSP JOURNAL
影响因子: --
作者:
Doya, Kenji
通讯作者: Doya, Kenji