Multi-layer network utilizing rewarded spike time dependent plasticity to learn a foraging task.

Multi-layer network utilizing rewarded spike time dependent plasticity to learn a foraging task.
复制标题

DOI:
10.1371/journal.pcbi.1005705
复制
发表时间:
2017-09
影响因子:
4.3
通讯作者:
Bazhenov M
Bazhenov M
中科院分区:
生物学2区
文献类型:
--
作者:
Sanda P;Skorheim S;Bazhenov M

文献摘要

参考文献

被引文献

相似文献

具有单一塑性层的神经网络采用奖励调制尖峰时间依赖可塑性(STDP)能够学习简单的觅食任务。在这里,我们展示了先进的模式识别和连续学习的尖峰神经元与多个塑料层的网络。该网络利用奖励调制和非奖励调制的STDP,并实现了多种机制,用于突触功效的稳态调节,包括异突触可塑性,增益控制,输出平衡,奖励STDP的活动正常化和突触强度的硬限制。我们发现,增加一个隐藏层的神经元采用非奖励STDP创建神经元,响应特定的输入组合,从而执行基本的分类的输入模式。当与下面一层实现奖励STDP的神经元相结合时,尽管没有标记的训练数据,该网络仍然能够学习区分奖励模式和指定为惩罚的模式。突触噪音允许试错学习,有助于确定有效的任务解决的目标导向的策略。该研究预测了具有STDP的尖峰神经元网络的一组关键属性,这些属性足以解决涉及模式分类和决策的复杂觅食任务。这项研究探讨了智能行为是如何从生物神经网络动力学的细胞水平上的基本原理中出现的。与人工智能社区中使用的方法相比,我们应用了神经元动力学和可塑性的生物逼真建模。该模型的构建模块是尖峰神经元,尖峰时间依赖可塑性(STDP)和稳态规则,实验已知,这被证明在保持网络稳定和能够连续学习方面发挥了重要作用。我们的研究预测,这些原则的组合使得可能的觅食行为在一个以前未知的环境中,包括模式分类,以区分奖励和惩罚的环境形状和决策,以选择最佳策略,以获得最大数量的奖励元素。为了解决这个复杂的任务,我们使用了多层神经元处理,在早期处理步骤中通过无监督STDP实现模式泛化,正如在动物和人类感觉处理中通常观察到的那样,然后在后期步骤中进行强化学习。在该模型中,智能行为自发地出现,由于网络组织实现了局部无监督可塑性和奖励反馈,导致在环境中的成功行为。
Neural networks with a single plastic layer employing reward modulated spike time dependent plasticity (STDP) are capable of learning simple foraging tasks. Here we demonstrate advanced pattern discrimination and continuous learning in a network of spiking neurons with multiple plastic layers. The network utilized both reward modulated and non-reward modulated STDP and implemented multiple mechanisms for homeostatic regulation of synaptic efficacy, including heterosynaptic plasticity, gain control, output balancing, activity normalization of rewarded STDP and hard limits on synaptic strength. We found that addition of a hidden layer of neurons employing non-rewarded STDP created neurons that responded to the specific combinations of inputs and thus performed basic classification of the input patterns. When combined with a following layer of neurons implementing rewarded STDP, the network was able to learn, despite the absence of labeled training data, discrimination between rewarding patterns and the patterns designated as punishing. Synaptic noise allowed for trial-and-error learning that helped to identify the goal-oriented strategies which were effective in task solving. The study predicts a critical set of properties of the spiking neuronal network with STDP that was sufficient to solve a complex foraging task involving pattern classification and decision making. This study explores how intelligent behavior emerges from the basic principles known at the cellular level of biological neuronal network dynamics. Compared to the approaches used in the artificial intelligence community, we applied biologically realistic modeling of neuronal dynamics and plasticity. The building blocks of the model are spiking neurons, spike-time dependent plasticity (STDP) and homeostatic rules, known experimentally, which are shown to play a fundamental role in both keeping the network stable and capable of continous learning. Our study predicts that a combination of these principles makes possible a foraging behavior in a previously unknown environment, including pattern classification to distinct between environment shapes which are rewarded and those which are punished and decision making to select the optimal strategy to acquire the maximal number of the rewarded elements. To solve this complex task we used multi-layer neuronal processing that implemented pattern generalization by unsupervised STDP at the earlier processing step, as commonly observed in the animal and human sensory processing, followed by reinforcement learning at the later steps. In the model, the intelligent behavior emerged spontaneously due to the network organization implementing both local unsupervised plasticity and reward feedback resulting from a successful behavior in the environment.
DOI: 10.3389/fncom.2015.00089
发表时间: 2015
影响因子: 3.2
作者:
Chistiakova M;Bannon NM;Chen JY;Bazhenov M;Volgushev M
通讯作者: Volgushev M
新皮质的异质突触可塑性。
DOI: 10.1007/s00221-009-1859-5
发表时间: 2009-12
影响因子: 2
作者:
Chistiakova, Marina;Volgushev, Maxim
通讯作者: Volgushev, Maxim
DOI: 10.1038/nn.2479
发表时间: 2010-03-01
影响因子: 25
作者:
Clopath, Claudia;Buesing, Lars;Gerstner, Wulfram
通讯作者: Gerstner, Wulfram
DOI: 10.1038/365337a0
发表时间: 1993-09-23
期刊: NATURE
影响因子: 64.8
作者:
DOUGLASS, JK;WILKENS, L;MOSS, F
通讯作者: MOSS, F
DOI: 10.1016/j.neuron.2010.07.023
发表时间: 2010-08-12
期刊: Neuron
影响因子: 16.2
作者:
Bazhenov M;Stopfer M
通讯作者: Stopfer M