Oculomotor learning revisited: a model of reinforcement learning in the basal ganglia incorporating an efference copy of motor actions.

Oculomotor learning revisited: a model of reinforcement learning in the basal ganglia incorporating an efference copy of motor actions.
复制标题

DOI:
10.3389/fncir.2012.00038
复制
发表时间:
2012
影响因子:
3.5
通讯作者:
Fee MS
Fee MS
中科院分区:
医学3区
文献类型:
--
作者:
Fee MS

文献摘要

参考文献

被引文献

相似文献

强化学习最简单的表述是,如果在特定环境中采取的行动之后有一个有利的结果,那么在相同的环境中,产生该行动的趋势应该得到加强或加强。虽然强化学习形成了许多当前基底神经节(BG)功能理论的基础,但这些模型并没有为传达上下文的信号和传达动物采取什么行动的信号整合不同的计算角色。最近的实验表明,鸣禽的声音相关的BG电路接收两个功能不同的兴奋性输入。一个输入来自皮层区域,该区域携带关于运动序列中的当前“时间”的上下文信息。另一个是来自一个单独的大脑皮层区域的运动命令的传出副本,该区域在学习过程中产生声音变化。基于这些发现,我在这里提出了一个一般模型的脊椎动物BG功能,结合上下文信息与一个独特的运动传出复制信号。这些信号通过学习规则进行整合,其中传出复制输入门控上下文输入(但不是传出复制输入)的增强作用到中等多刺神经元上以响应奖励动作。该假设被描述在一个电路,实现视觉引导的眼跳的学习。该模型使可测试的预测的解剖和功能特性的假设的上下文和传出复制输入到纹状体从丘脑和皮质来源。
In its simplest formulation, reinforcement learning is based on the idea that if an action taken in a particular context is followed by a favorable outcome, then, in the same context, the tendency to produce that action should be strengthened, or reinforced. While reinforcement learning forms the basis of many current theories of basal ganglia (BG) function, these models do not incorporate distinct computational roles for signals that convey context, and those that convey what action an animal takes. Recent experiments in the songbird suggest that vocal-related BG circuitry receives two functionally distinct excitatory inputs. One input is from a cortical region that carries context information about the current “time” in the motor sequence. The other is an efference copy of motor commands from a separate cortical brain region that generates vocal variability during learning. Based on these findings, I propose here a general model of vertebrate BG function that combines context information with a distinct motor efference copy signal. The signals are integrated by a learning rule in which efference copy inputs gate the potentiation of context inputs (but not efference copy inputs) onto medium spiny neurons in response to a rewarded action. The hypothesis is described in terms of a circuit that implements the learning of visually guided saccades. The model makes testable predictions about the anatomical and functional properties of hypothesized context and efference copy inputs to the striatum from both thalamic and cortical sources.
神经元型特异性信号,用于腹侧对段区域的奖励和惩罚。
DOI: 10.1038/nature10754
发表时间: 2012-01-18
期刊: NATURE
影响因子: 64.8
作者:
Cohen, Jeremiah Y.;Haesler, Sebastian;Vong, Linh;Lowell, Bradford B.;Uchida, Naoshige
通讯作者: Uchida, Naoshige
DOI: 10.1073/pnas.94.13.7036
发表时间: 1997-06-24
影响因子: 11.1
作者:
Charpier, S;Deniau, JM
通讯作者: Deniau, JM
DOI: 10.1126/science.1155140
发表时间: 2008-05-02
期刊: SCIENCE
影响因子: 56.9
作者:
Aronov, Dmitriy;Andalman, Aaron S.;Fee, Michale S.
通讯作者: Fee, Michale S.
DOI: 10.1002/cne.902710308
发表时间: 1988-05-15
影响因子: 2.5
作者:
ABRAMSON, BP;CHALUPA, LM
通讯作者: CHALUPA, LM
DOI: 10.1523/jneurosci.22-05-01883.2002
发表时间: 2002-03-01
影响因子: 5.3
作者:
Basso, MA;Wurtz, RH
通讯作者: Wurtz, RH