Reinforcement learning, spike-time-dependent plasticity, and the BCM rule

Reinforcement learning, spike-time-dependent plasticity, and the BCM rule
复制标题

DOI:
10.1162/neco.2007.19.8.2245
复制
发表时间:
2007-08-01
期刊:
影响因子:
2.9
通讯作者:
Meir, Ron
Meir, Ron
中科院分区:
计算机科学4区
文献类型:
--
作者:
Baras, Dorit;Meir, Ron

文献摘要

被引文献

相似文献

学习代理,无论是自然的还是人工的,都必须更新它们的内部参数,以便随着时间的推移改善它们的行为。在强化学习中,这种可塑性受到环境信号的影响,称为奖励,它将变化引导到适当的方向。我们应用最近推出的机器学习的策略学习算法的尖峰神经元网络,并得出一个尖峰时间依赖的可塑性规则,确保收敛到局部最优的预期平均奖励。该方法适用于广泛的一类神经元模型,包括Hodgkin-Huxley模型。我们证明了衍生规则在几个玩具问题的有效性。最后,通过统计分析,我们表明,建立的突触可塑性规则是密切相关的广泛使用的神经元突触可塑性规则,有很好的生物学证据。
Learning agents, whether natural or artificial, must update their internal parameters in order to improve their behavior over time. In reinforcement learning, this plasticity is influenced by an environmental signal, termed a reward, that directs the changes in appropriate directions. We apply a recently introduced policy learning algorithm from machine learning to networks of spiking neurons and derive a spike-time-dependent plasticity rule that ensures convergence to a local optimum of the expected average reward. The approach is applicable to a broad class of neuronal models, including the Hodgkin-Huxley model. We demonstrate the effectiveness of the derived rule in several toy problems. Finally, through statistical analysis, we show that the synaptic plasticity rule established is closely related to the widely used BCM rule, for which good biological evidence exists.