Reinforcement learning of targeted movement in a spiking neuronal model of motor cortex.

Reinforcement learning of targeted movement in a spiking neuronal model of motor cortex.
复制标题

DOI:
10.1371/journal.pone.0047251
复制
发表时间:
2012
期刊:
影响因子:
3.7
通讯作者:
Lytton WW
Lytton WW
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Chadderdon GL;Neymotin SA;Kerr CC;Lytton WW

文献摘要

参考文献

被引文献

相似文献

传统上,感觉运动控制是从控制理论的角度来考虑的,与神经生物学无关。相比之下,在这里我们利用运动皮层的尖峰神经元模型并训练它执行简单的运动任务,其中包括将单关节“前臂”旋转到目标。学习基于类似于多巴胺系统的强化机制。这分别提供了全局奖励或惩罚信号,以响应手到目标距离的减小或增加。输出部分由泊松运动驱动,产生随机运动,然后可以通过学习来塑造。虚拟前臂由围绕肘关节旋转的单个节段组成,由屈肌和伸肌控制。该模型由 144 个兴奋性神经元和 64 个基于事件的抑制性神经元组成,每个神经元都具有 AMPA、NMDA 和 GABA 突触。该模型的本体感受细胞输入编码了 2 个肌肉长度。可塑性仅在输入和输出兴奋单元之间的前馈连接中启用,使用依赖于尖峰时序的资格轨迹进行突触信用或责备分配。学习由全局三值信号产生:奖励(+1)、无学习(0)或惩罚(−1),分别对应于多巴胺能细胞放电的阶段性增加、缺乏变化或阶段性减少。只有当奖励和惩罚同时启用时,成功的学习才会发生。在这种情况下,在 180 秒的模拟时间内成功学习了 5 个目标角度,中位误差为 8 度。电机咿呀学语允许探索性学习,但降低了学习行为的稳定性,因为手在达到目标后继续移动。我们的模型证明,全局强化信号与突触可塑性的资格痕迹相结合,可以训练尖峰感觉运动网络来执行目标导向的运动行为。
Sensorimotor control has traditionally been considered from a control theory perspective, without relation to neurobiology. In contrast, here we utilized a spiking-neuron model of motor cortex and trained it to perform a simple movement task, which consisted of rotating a single-joint “forearm” to a target. Learning was based on a reinforcement mechanism analogous to that of the dopamine system. This provided a global reward or punishment signal in response to decreasing or increasing distance from hand to target, respectively. Output was partially driven by Poisson motor babbling, creating stochastic movements that could then be shaped by learning. The virtual forearm consisted of a single segment rotated around an elbow joint, controlled by flexor and extensor muscles. The model consisted of 144 excitatory and 64 inhibitory event-based neurons, each with AMPA, NMDA, and GABA synapses. Proprioceptive cell input to this model encoded the 2 muscle lengths. Plasticity was only enabled in feedforward connections between input and output excitatory units, using spike-timing-dependent eligibility traces for synaptic credit or blame assignment. Learning resulted from a global 3-valued signal: reward (+1), no learning (0), or punishment (−1), corresponding to phasic increases, lack of change, or phasic decreases of dopaminergic cell firing, respectively. Successful learning only occurred when both reward and punishment were enabled. In this case, 5 target angles were learned successfully within 180 s of simulation time, with a median error of 8 degrees. Motor babbling allowed exploratory learning, but decreased the stability of the learned behavior, since the hand continued moving after reaching the target. Our model demonstrated that a global reinforcement signal, coupled with eligibility traces for synaptic plasticity, can train a spiking sensorimotor network to perform goal-directed motor behavior.
DOI: 10.1038/nature04766
发表时间: 2006-06-15
期刊: NATURE
影响因子: 64.8
作者:
Daw, Nathaniel D.;O'Doherty, John P.;Dayan, Peter;Seymour, Ben;Dolan, Raymond J.
通讯作者: Dolan, Raymond J.
DOI: 10.1152/jn.00364.2007
发表时间: 2007-12-01
影响因子: 2.5
作者:
Farries, Michael A.;Fairhall, Adrienne L.
通讯作者: Fairhall, Adrienne L.
DOI: 10.1093/cercor/8.4.346
发表时间: 1998-06-01
期刊: CEREBRAL CORTEX
影响因子: 3.7
作者:
Almássy, N;Edelman, GM;Sporns, O
通讯作者: Sporns, O
DOI: 10.1097/wnp.0b013e3180336fc0
发表时间: 2007-04-01
影响因子: 2.4
作者:
Lytton, William W.;Omurtag, Ahmet
通讯作者: Omurtag, Ahmet
DOI: 10.1093/cercor/bhl152
发表时间: 2007-10-01
期刊: CEREBRAL CORTEX
影响因子: 3.7
作者:
Izhikevich, Eugene M.
通讯作者: Izhikevich, Eugene M.