An Efficient Hardware Implementation of Reinforcement Learning: The Q-Learning Algorithm

An Efficient Hardware Implementation of Reinforcement Learning: The Q-Learning Algorithm
复制标题

DOI:
10.1109/access.2019.2961174
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Re, Marco
Re, Marco
中科院分区:
计算机科学3区
文献类型:
--
作者:
Spano, Sergio;Cardarilli, Gian Carlo;Re, Marco

文献摘要

被引文献

相似文献

在本文中,我们提出了一种有效的硬件体系结构,该架构实现了Q学习算法,适用于实时应用程序。它的主要功能是低功率,高吞吐量和有限的硬件资源。我们还提出了一种基于近似乘数的技术,以降低算法的硬件复杂性。我们在Xilinx Zynq Ultrascale + MPSOC ZCU106评估套件上实现了设计。实现结果是根据硬件资源,吞吐量和功耗评估的。将体系结构与文献中介绍的Q学习硬件加速器的艺术状态进行了比较,从而获得了更好的速度,功率和硬件资源的结果。提出了使用不同尺寸的Q-matrix和固定点算术的不同词根长度的实验。 Q-Matrix的尺寸为8 x 4(8位数据),我们达到了222个MSP(每秒巨型样品)的吞吐量和37 MW的动态功率消耗,而Q-Matrix的大小为256 x 16(32)位数据)我们达到了93个MSP的吞吐量和611 MW的功耗。由于加速器所需的硬件资源少,我们的系统适用于多代理IoT应用程序。此外,该体系结构可用于实现SARSA(州行动奖励 - 状态)的增强算法,并进行了较小的修改。
In this paper we propose an efficient hardware architecture that implements the Q-Learning algorithm, suitable for real-time applications. Its main features are low-power, high throughput and limited hardware resources. We also propose a technique based on approximated multipliers to reduce the hardware complexity of the algorithm. We implemented the design on a Xilinx Zynq Ultrascale + MPSoC ZCU106 Evaluation Kit. The implementation results are evaluated in terms of hardware resources, throughput and power consumption. The architecture is compared to the state of the art of Q-Learning hardware accelerators presented in the literature obtaining better results in speed, power and hardware resources. Experiments using different sizes for the Q-Matrix and different wordlengths for the fixed point arithmetic are presented. With a Q-Matrix of size 8 x 4 (8 bit data) we achieved a throughput of 222 MSPS (Mega Samples Per Second) and a dynamic power consumption of 37 mW, while with a Q-Matrix of size 256 x 16 (32 bit data) we achieved a throughput of 93 MSPS and a power consumption 611 mW. Due to the small amount of hardware resources required by the accelerator, our system is suitable for multi-agent IoT applications. Moreover, the architecture can be used to implement the SARSA (State-Action-Reward-State-Action) Reinforcement Learning algorithm with minor modifications.