Mixed-Precision Deep Learning Based on Computational Memory

Mixed-Precision Deep Learning Based on Computational Memory
复制标题

DOI:
10.3389/fnins.2020.00406
复制
发表时间:
2020-05-12
影响因子:
4.3
通讯作者:
Eleftheriou, Evangelos
Eleftheriou, Evangelos
中科院分区:
医学2区
文献类型:
--
作者:
Nandakumar, S. R.;Le Gallo, Manuel;Eleftheriou, Evangelos

文献摘要

被引文献

相似文献

深度神经网络(DNN)彻底改变了人工智能领域,并在图像和语音识别等认知任务中取得了前所未有的成功。然而,大型DNN的训练是计算密集型的,这促使人们寻找针对该应用的新型计算架构。具有以交叉杆阵列组织的纳米级电阻存储器设备的计算存储器单元可以将突触权重存储在其电导状态中,并且以非冯·诺依曼方式在适当的位置执行昂贵的加权求和。然而,在权重更新过程期间以可靠的方式更新电导状态是限制这种实现的训练准确性的基本挑战。在这里,我们提出了一种混合精度的架构,它结合了一个计算存储器单元执行加权求和和不精确的电导更新与数字处理单元,积累了高精度的权重更新。基于所提出的架构,使用相变存储器(PCM)阵列的多层感知器的硬件/软件相结合的训练实验实现了97.73%的测试准确率的任务分类手写数字(基于MNIST数据集),在0.6%的软件基线。该架构使用PCM的精确行为模型在广泛的一类网络上进行进一步评估,即卷积神经网络,长短期记忆网络和生成对抗网络。精度可比的浮点实现,而不受约束的非理想与PCM设备。系统级研究表明,与专用的全数字32位实现相比,用于训练多层感知器时,该架构的能效提高了172倍。
Deep neural networks (DNNs) have revolutionized the field of artificial intelligence and have achieved unprecedented success in cognitive tasks such as image and speech recognition. Training of large DNNs, however, is computationally intensive and this has motivated the search for novel computing architectures targeting this application. A computational memory unit with nanoscale resistive memory devices organized in crossbar arrays could store the synaptic weights in their conductance states and perform the expensive weighted summations in place in a non-von Neumann manner. However, updating the conductance states in a reliable manner during the weight update process is a fundamental challenge that limits the training accuracy of such an implementation. Here, we propose a mixed-precision architecture that combines a computational memory unit performing the weighted summations and imprecise conductance updates with a digital processing unit that accumulates the weight updates in high precision. A combined hardware/software training experiment of a multilayer perceptron based on the proposed architecture using a phase-change memory (PCM) array achieves 97.73% test accuracy on the task of classifying handwritten digits (based on the MNIST dataset), within 0.6% of the software baseline. The architecture is further evaluated using accurate behavioral models of PCM on a wide class of networks, namely convolutional neural networks, long-short-term-memory networks, and generative-adversarial networks. Accuracies comparable to those of floating-point implementations are achieved without being constrained by the non-idealities associated with the PCM devices. A system-level study demonstrates 172 x improvement in energy efficiency of the architecture when used for training a multilayer perceptron compared with a dedicated fully digital 32-bit implementation.