Functional Requirements for Reward-Modulated Spike-Timing-Dependent Plasticity

Functional Requirements for Reward-Modulated Spike-Timing-Dependent Plasticity
复制标题

DOI:
10.1523/jneurosci.6249-09.2010
复制
发表时间:
2010-10-06
影响因子:
5.3
通讯作者:
Gerstner, Wulfram
Gerstner, Wulfram
中科院分区:
医学1区
文献类型:
--
作者:
Fremaux, Nicolas;Sprekeler, Henning;Gerstner, Wulfram

文献摘要

被引文献

相似文献

最近的实验表明,尖峰时间依赖的可塑性受到神经调制的影响。我们推导出成功学习的奖励相关的行为的理论条件为一大类的学习规则,其中赫布突触可塑性的条件是一个全球性的调制因子信号奖励。我们表明,在这个类中的所有学习规则可以分为一个术语,捕捉神经元放电和奖励的协方差和第二个术语,提出了无监督学习的影响。无监督的长期,这是,在一般情况下,有害的奖励为基础的学习,可以被抑制,如果神经调节信号编码之间的差异奖励和预期的奖励,但只有当预期的奖励是计算每个任务和刺激分开。如果要同时学习多项任务,神经系统就需要一个能够预测任意刺激的预期奖励的内部批评家。我们表明,与批评家,奖励调制尖峰时间依赖性可塑性是能够学习运动轨迹的时间分辨率为几十毫秒。时间差学习的关系,块为基础的学习范式的相关性,并与批评学习的局限性进行了讨论。
Recent experiments have shown that spike-timing-dependent plasticity is influenced by neuromodulation. We derive theoretical conditions for successful learning of reward-related behavior for a large class of learning rules where Hebbian synaptic plasticity is conditioned on a global modulatory factor signaling reward. We show that all learning rules in this class can be separated into a term that captures the covariance of neuronal firing and reward and a second term that presents the influence of unsupervised learning. The unsupervised term, which is, in general, detrimental for reward-based learning, can be suppressed if the neuromodulatory signal encodes the difference between the reward and the expected reward-but only if the expected reward is calculated for each task and stimulus separately. If several tasks are to be learned simultaneously, the nervous system needs an internal critic that is able to predict the expected reward for arbitrary stimuli. We show that, with a critic, reward-modulated spike-timing-dependent plasticity is capable of learning motor trajectories with a temporal resolution of tens of milliseconds. The relation to temporal difference learning, the relevance of block-based learning paradigms, and the limitations of learning with a critic are discussed.