Combining Supervised, Unsupervised, and Reinforcement Learning in a Network of Spiking Neurons

Combining Supervised, Unsupervised, and Reinforcement Learning in a Network of Spiking Neurons
复制标题

DOI:
10.1007/978-90-481-9695-1_26
复制
发表时间:
2011
期刊:
2018 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Sebastian Handrich;A. Herzog;A. Wolf;C. Herrmann
Sebastian Handrich;A. Herzog;A. Wolf;C. Herrmann
中科院分区:
其他
文献类型:
--
作者:
Sebastian Handrich;A. Herzog;A. Wolf;C. Herrmann

文献摘要

被引文献

相似文献

人类大脑通过多种不同的学习策略不断学习。它可以通过简单地向其感觉器官提供刺激来学习,这被认为是无监督学习。此外,当教师提供输出时,它可以学习输入和输出之间的关联,这被认为是监督学习。最重要的是,如果正确的行为之后是奖励和/或不正确的行为之后是惩罚,它可以非常有效地学习,这被认为是强化学习。到目前为止,大多数人工神经架构只实现了三种学习机制中的一种,尽管大脑集成了所有三种学习机制。在这里,我们在尖峰神经元网络中实现了无监督,监督和强化学习。为了实现这一雄心勃勃的目标,现有的学习规则称为尖峰时间依赖可塑性,必须扩展,使其受到奖励信号多巴胺的调制。
The human brain constantly learns via mutiple different learning strategies. It can learn by simply having stimuli being presented to its sensory organs which is considered unsupervised learning. In addition, it can learn associations between inputs and outputs when a teacher provides the output which is considered as supervised learning. Most importantly, it can learn very efficiently if correct behaviour is followed by reward and/or incorrect behaviour is followed by punishment which is considered reinforcement learning. So far, most artificial neural architectures implement only one of the three learning mechanisms — even though the brain integrates all three. Here, we have implemented unsupervised, supervised, and reinforcement learning within a network of spiking neurons. In order to achieve this ambitious goal, the existing learning rule called spike-timing-dependent plasticity had to be extended such that it is modulated by the reward signal dopamine.