Learning in spiking neural networks by reinforcement of stochastic synaptic transmission

Learning in spiking neural networks by reinforcement of stochastic synaptic transmission
复制标题

DOI:
10.1016/s0896-6273(03)00761-x
复制
发表时间:
2003-12-18
期刊:
影响因子:
16.2
通讯作者:
Seung, HS
Seung, HS
中科院分区:
医学1区
文献类型:
--
作者:
Seung, HS

文献摘要

被引文献

相似文献

众所周知,化学突触传递是一个不可靠的过程,但这种不可靠性的功能仍然不清楚。在这里,我考虑的假设是,突触传递的随机性被大脑利用来学习,类似于达尔文进化论利用基因突变的方式。如果突触是“享乐主义”的,这是可能的,它通过增加囊泡释放或失败的概率来响应全局奖励信号,这取决于奖励之前立即采取的行动。快乐突触通过计算平均奖励梯度的随机近似来学习。它们与突触动力学(如短期易化和抑制)以及树突整合和动作电位产生的复杂性相容。通过适当地给予奖励,可以训练享乐主义突触网络来执行所需的计算,如这里通过积分和激发模型神经元的数值模拟所示。
It is well-known that chemical synaptic transmission is an unreliable process, but the function of such unreliability remains unclear. Here I consider the hypothesis that the randomness of synaptic transmission is harnessed by the brain for learning, in analogy to the way that genetic mutation is utilized by Darwinian evolution. This is possible if synapses are "hedonistic," responding to a global reward signal by increasing their probabilities of vesicle release or failure, depending on which action immediately preceded reward. Hedonistic synapses learn by computing a stochastic approximation to the gradient of the average reward. They are compatible with synaptic dynamics such as short-term facilitation and depression and with the intricacies of dendritic integration and action potential generation. A network of hedonistic synapses can be trained to perform a desired computation by administering reward appropriately, as illustrated here through numerical simulations of integrate-and-fire model neurons.