Hardware-Friendly Actor-Critic Reinforcement Learning Through Modulation of Spike-Timing-Dependent Plasticity

Hardware-Friendly Actor-Critic Reinforcement Learning Through Modulation of Spike-Timing-Dependent Plasticity
复制标题

DOI:
10.1109/tc.2016.2595580
复制
发表时间:
2017-02
影响因子:
3.7
通讯作者:
Nan Zheng;P. Mazumder
Nan Zheng;P. Mazumder
中科院分区:
计算机科学2区
文献类型:
--
作者:
Nan Zheng;P. Mazumder

文献摘要

被引文献

相似文献

在这项工作中,我们提出了一种硬件友好的强化学习算法。该学习算法基于用尖峰神经网络(SNN)实现的行动者-批评者结构。提出了一种生物上合理的、硬件友好的、依赖于脉冲定时的可塑性学习规则,并将其用于SNN的训练。研究了在强化学习环境中应用学习规则的几个重要方面,特别是从电路设计者的角度。潜在的噪声混合和相关尖峰的陷阱被识别和适当地处理。为了具有低功耗的学习结构,提出了诸如对特定学习块的数据进行下采样、将量化噪声作为噪声残留物注入神经元以及适当的记忆划分等技术。文中以一维状态值函数学习问题和二维迷宫行走问题为例,说明了所提算法和学习规则的有效性。提出了一种低功耗的硬件结构,并用Verilog实现了实例。分析了该算法的硬件复杂性,并讨论了当问题规模变大时打破内存瓶颈的可能解决方案。
In this work, we propose a hardware-friendly reinforcement learning algorithm. The learning algorithm is based on an actor-critic structure implemented with spiking neural networks (SNNs). A biologically plausible and hardware-friendly spike-timing-dependent plasticity learning rule is formulated and employed in the training of SNNs. Several important aspects of applying the learning rule in a reinforcement learning context is studied, especially from the circuit designers’ point of view. Pitfalls of potential noise mixing and correlated spikes are identified and properly addressed. To feature a low-power learning architecture, techniques such as down-sampling data for certain learning blocks, injecting quantization noise as noisy residues in neurons, and proper memory partitioning are proposed. A 1-D state-value function learning problem and a 2-D maze walking problem are examined in this paper to illustrate effectiveness of the proposed algorithm and learning rules. A low-power hardware architecture is proposed and examples are implemented with Verilog. Hardware complexity of the proposed algorithm is analyzed, and potential solutions to breaking memory bottleneck when the size of the problem gets large is also discussed.