TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks

TIME: A Training-in-Memory Architecture for RRAM-Based Deep Neural Networks
复制标题

DOI:
10.1109/tcad.2018.2824304
复制
发表时间:
2019-05
影响因子:
2.9
通讯作者:
Ming Cheng;Lixue Xia;Zhenhua Zhu;Yi Cai;Yuan Xie;Yu Wang;Huazhong Yang
Ming Cheng;Lixue Xia;Zhenhua Zhu;Yi Cai;Yuan Xie;Yu Wang;Huazhong Yang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Ming Cheng;Lixue Xia;Zhenhua Zhu;Yi Cai;Yuan Xie;Yu Wang;Huazhong Yang

文献摘要

被引文献

相似文献

神经网络的训练通常耗时且耗费大量资源。新兴的金属氧化物电阻式随机存取存储器(RRAM)显示出神经网络计算的潜力。RRAM的横条结构和多比特特性可以高效地进行矩阵向量积,这是神经网络最常见的运算。实现基于随机存储器的神经网络训练存在两个挑战。首先,目前基于RRAM的结构只支持训练神经网络中的推理,不能进行训练神经网络的反向传播(BP)和权值更新。其次,训练NN需要大量的迭代来不断更新权值以达到收敛。然而,由于RRAM的非理想因素,这种权重更新导致了大量的能量消耗。在本文中,我们提出了一种基于RRAM (TIME)架构的内存中训练方法,并设计了外围电路来实现在RRAM上训练神经网络。TIME支持BP和权值更新,同时最大限度地重用RRAM上推理操作的外围电路。同时,设计了一套针对非理想因素的优化策略,以降低RRAM的优化成本。我们探讨了监督学习(SL)和深度强化学习(DRL)在时间上的性能。为了进一步提高能源效率,本文还介绍了一种具体的DRL映射方法。仿真结果表明,与采用CMOS技术的专用集成电路(ASIC) DaDianNao相比,TIME在SL中的平均能效可提高5.3{\times}$。在DRL中,TIME的能效平均比GPU高126{\ \}$。如果可以进一步降低RRAM的调整成本,TIME有可能将能源效率提高两个数量级,而不是ASIC。
The training of neural networks (NN) is usually time-consuming and resource intensive. The emerging metal-oxide resistive random-access memory (RRAM) device has shown potential for the computation of NN. RRAM crossbar structure and multibit characteristics can perform the matrix-vector product in high energy efficiency, which is the most common operation of NN. Two challenges exist for realizing training NN based on RRAM. First, the current architectures based on RRAM only support the inference in training NN and cannot perform the backpropagation (BP) and the weight update of training NN. Second, training NN requires enormous iterations to constantly update the weights for reaching the convergence. However, this weight update leads to large energy consumption because of the nonideal factors of RRAM. In this paper, we propose a training-in-memory based on RRAM (TIME) architecture and the peripheral circuit design to enable training NN on RRAM. TIME supports the BP and the weight update while maximizing the re-usage of peripheral circuits of the inference operation on RRAM. Meanwhile, a set of optimization strategies focusing on the nonideal factors are designed to reduce the cost of tuning RRAM. We explore the performance of both supervised learning (SL) and deep reinforcement learning (DRL) on TIME. A specific mapping method of DRL is also introduced to further improve energy efficiency. Simulation results show that in SL, TIME can achieve $5.3{\times }$ higher energy efficiency on average compared with DaDianNao, an application-specific integrated circuits (ASIC) in CMOS technology. In DRL, TIME can perform an average $126{\times }$ higher than GPU in energy efficiency. If the cost of tuning RRAM can be further reduced, TIME has the potential to boost the energy efficiency by two orders of magnitudes compared with ASIC.