Multi-layer network utilizing rewarded spike time dependent plasticity to learn a foraging task.
Multi-layer network utilizing rewarded spike time dependent plasticity to learn a foraging task.
复制标题
DOI:
10.1371/journal.pcbi.1005705
复制
发表时间:
2017-09
影响因子:
4.3
通讯作者:
Bazhenov M
中科院分区:
文献类型:
--
作者:
Sanda P;Skorheim S;Bazhenov M
Neural networks with a single plastic layer employing reward modulated spike time dependent plasticity (STDP) are capable of learning simple foraging tasks. Here we demonstrate advanced pattern discrimination and continuous learning in a network of spiking neurons with multiple plastic layers. The network utilized both reward modulated and non-reward modulated STDP and implemented multiple mechanisms for homeostatic regulation of synaptic efficacy, including heterosynaptic plasticity, gain control, output balancing, activity normalization of rewarded STDP and hard limits on synaptic strength. We found that addition of a hidden layer of neurons employing non-rewarded STDP created neurons that responded to the specific combinations of inputs and thus performed basic classification of the input patterns. When combined with a following layer of neurons implementing rewarded STDP, the network was able to learn, despite the absence of labeled training data, discrimination between rewarding patterns and the patterns designated as punishing. Synaptic noise allowed for trial-and-error learning that helped to identify the goal-oriented strategies which were effective in task solving. The study predicts a critical set of properties of the spiking neuronal network with STDP that was sufficient to solve a complex foraging task involving pattern classification and decision making. This study explores how intelligent behavior emerges from the basic principles known at the cellular level of biological neuronal network dynamics. Compared to the approaches used in the artificial intelligence community, we applied biologically realistic modeling of neuronal dynamics and plasticity. The building blocks of the model are spiking neurons, spike-time dependent plasticity (STDP) and homeostatic rules, known experimentally, which are shown to play a fundamental role in both keeping the network stable and capable of continous learning. Our study predicts that a combination of these principles makes possible a foraging behavior in a previously unknown environment, including pattern classification to distinct between environment shapes which are rewarded and those which are punished and decision making to select the optimal strategy to acquire the maximal number of the rewarded elements. To solve this complex task we used multi-layer neuronal processing that implemented pattern generalization by unsupervised STDP at the earlier processing step, as commonly observed in the animal and human sensory processing, followed by reinforcement learning at the later steps. In the model, the intelligent behavior emerged spontaneously due to the network organization implementing both local unsupervised plasticity and reward feedback resulting from a successful behavior in the environment.
登录
查看更多内容
影响因子:
3.2
作者:
Chistiakova M;Bannon NM;Chen JY;Bazhenov M;Volgushev M
通讯作者:
Volgushev M
影响因子:
2
作者:
Chistiakova, Marina;Volgushev, Maxim
通讯作者:
Volgushev, Maxim
影响因子:
25
作者:
Clopath, Claudia;Buesing, Lars;Gerstner, Wulfram
通讯作者:
Gerstner, Wulfram
影响因子:
64.8
作者:
DOUGLASS, JK;WILKENS, L;MOSS, F
通讯作者:
MOSS, F
影响因子:
16.2
作者:
Bazhenov M;Stopfer M
通讯作者:
Stopfer M