A 55-nm, 1.0–0.4V, 1.25-pJ/MAC Time-Domain Mixed-Signal Neuromorphic Accelerator With Stochastic Synapses for Reinforcement Learning in Autonomous Mobile Robots

A 55-nm, 1.0–0.4V, 1.25-pJ/MAC Time-Domain Mixed-Signal Neuromorphic Accelerator With Stochastic Synapses for Reinforcement Learning in Autonomous Mobile Robots
复制标题

具有随机突触的 55 nm、1.0–0.4V、1.25 pJ/MAC 时域混合信号神经形态加速器,用于自主移动机器人的强化学习

DOI:
10.1109/jssc.2018.2881288
复制
发表时间:
2019
影响因子:
5.4
通讯作者:
A. Raychowdhury
A. Raychowdhury
中科院分区:
工程技术1区
文献类型:
--
作者:
Anvesha Amaravati;Saad Bin Nasir;Justin Ting;Insik Yoon;A. Raychowdhury

文献摘要

被引文献

相似文献

强化学习(RL)是一种仿生学习方法,其中智能体可以在没有任何人类监督的情况下通过执行特定任务来了解环境。强化学习受到行为心理学的启发,在行为心理学中,代理人采取行动以最大化累积奖励。在本文中,我们提出了一个RL神经形态加速器,能够在云边缘的移动机器人中执行避障。我们提出了一种节能的时域混合信号(TD-MS)计算框架。在TD-MS计算中,我们证明了计算的能量与计算的重要性成正比。我们利用随机网络的独特特性和q学习的最新进展,在提议的强化学习实现中。55nm测试芯片采用三层全连接神经网络实现RL,峰值功耗为690 $\mu \text{W}$。
Reinforcement learning (RL) is a bio-mimetic learning approach, where agents can learn about an environment by performing specific tasks without any human supervision. RL is inspired by behavioral psychology, where agents take actions to maximize a cumulative reward. In this paper, we present an RL neuromorphic accelerator capable of performing obstacle avoidance in a mobile robot at the edge of the cloud. We propose an energy-efficient time-domain mixed-signal (TD-MS) computational framework. In TD-MS computation, we demonstrate that the energy to compute is proportional to the importance of the computation. We leverage the unique properties of stochastic networks and recent advances in Q-learning in the proposed RL implementation. The 55-nm test chip implements RL using a three-layered fully connected neural network and consumes a peak power of 690 $\mu \text{W}$ .