Equivalent-accuracy accelerated neural-network training using analogue memory

Equivalent-accuracy accelerated neural-network training using analogue memory
复制标题

DOI:
10.1038/s41586-018-0180-5
复制
发表时间:
2018-06-07
期刊:
影响因子:
64.8
通讯作者:
Burr, Geoffrey W.
Burr, Geoffrey W.
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Ambrogio, Stefano;Narayanan, Pritish;Burr, Geoffrey W.

文献摘要

被引文献

相似文献

神经网络训练可能会很缓慢且能耗高,这是因为需要在传统数字存储芯片和处理器芯片之间传输网络的权重数据。模拟非易失性存储器可以通过在权重数据所在位置的模拟域中执行并行化的乘积累加运算来加速被称为反向传播的神经网络训练算法。然而,由于动态范围不足和权重更新不对称性过大,这种使用非易失性存储器硬件的原位训练的分类准确率通常低于基于软件的训练。在此我们展示了混合的软硬件神经网络实现方式,其涉及多达204,900个突触,并且结合了相变存储器中的长期存储、易失性电容器的近线性更新以及带有“极性反转”的权重数据传输,以抵消固有的器件间差异。我们在各种常用的机器学习测试数据集(MNIST、MNIST - 背景随机化、CIFAR - 10和CIFAR - 100)上实现了与基于软件的训练相当的泛化准确率(针对之前未见过的数据)。我们为我们的实现所计算出的每秒每瓦280.65亿次运算的计算能效以及每平方毫米每秒3.6万亿次运算的单位面积吞吐量比当今的图形处理单元高出两个数量级。这项工作为实现既快速又节能的硬件加速器提供了一条途径,特别是在全连接神经网络层上。
Neural-network training can be slow and energy intensive, owing to the need to transfer the weight data for the network between conventional digital memory chips and processor chips. Analogue non-volatile memory can accelerate the neural-network training algorithm known as backpropagation by performing parallelized multiply-accumulate operations in the analogue domain at the location of the weight data. However, the classification accuracies of such in situ training using non-volatile-memory hardware have generally been less than those of software-based training, owing to insufficient dynamic range and excessive weight-update asymmetry. Here we demonstrate mixed hardware-software neural-network implementations that involve up to 204,900 synapses and that combine long-term storage in phase-change memory, near-linear updates of volatile capacitors and weight-data transfer with 'polarity inversion' to cancel out inherent device-to-device variations. We achieve generalization accuracies (on previously unseen data) equivalent to those of software-based training on various commonly used machine-learning test datasets (MNIST, MNIST-backrand, CIFAR-10 and CIFAR-100). The computational energy efficiency of 28,065 billion operations per second per watt and throughput per area of 3.6 trillion operations per second per square millimetre that we calculate for our implementation exceed those of today's graphical processing units by two orders of magnitude. This work provides a path towards hardware accelerators that are both fast and energy efficient, particularly on fully connected neural-network layers.