A CGRA based Neural Network Inference Engine for Deep Reinforcement Learning

A CGRA based Neural Network Inference Engine for Deep Reinforcement Learning
复制标题

DOI:
10.1109/apccas.2018.8605639
复制
发表时间:
2018-10
期刊:
2018 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS)
影响因子:
--
通讯作者:
Minglan Liang;Mingsong Chen;Zheng Wang;Jingwei Sun
Minglan Liang;Mingsong Chen;Zheng Wang;Jingwei Sun
中科院分区:
其他
文献类型:
--
作者:
Minglan Liang;Mingsong Chen;Zheng Wang;Jingwei Sun

文献摘要

被引文献

相似文献

最近人工智能算法的超快速发展需要专用的神经网络加速器,其高计算性能和低功耗使深度学习算法能够部署在边缘计算节点上。最先进的深度学习引擎大多支持监督学习,如CNN,RNN,而很少有AI引擎支持片上强化学习,这是自治系统决策子系统的最重要算法核心。在这项工作中,一个粗粒度的可重构阵列(CGRA)像AI计算引擎已被设计用于监督和强化学习的部署。基于65nmCMOS工艺在200MHz设计频率下的逻辑综合结果表明,该引擎的物理性能指标为:硅面积0.32mm2,功耗15.45mW。所提出的片上AI引擎促进了端到端感知和决策网络的实现,这些网络可以在自动驾驶、机器人和无人机中广泛应用。
Recent ultra-fast development of artificial intelligence algorithms has demanded dedicated neural network accelerators, whose high computing performance and low power consumption enable the deployment of deep learning algorithms on the edge computing nodes. State-of-the-art deep learning engines mostly support supervised learning such as CNN, RNN, whereas very few AI engines support on-chip reinforcement learning, which is the foremost algorithm kernel for decision-making subsystem of an autonomous system. In this work, a Coarse-grained Reconfigurable Array (CGRA) like AI computing engine has been designed for the deployments of both supervised and reinforcement learning. Logic synthesis at the design frequency of 200MHz based on 65nm CMOS technology reveals the physical statistics of the proposed engine of 0.32mm2 in silicon area, 15.45 mW in power consumption. The proposed on-chip AI engine facilitates the implementation of end-to-end perceptual and decision-making networks, which can find its wide employment in autonomous driving, robotics and UAVs.