Gradient Descent Using Stochastic Circuits for Efficient Training of Learning Machines

Gradient Descent Using Stochastic Circuits for Efficient Training of Learning Machines
复制标题

使用随机电路的梯度下降来有效训练学习机

DOI:
--
复制
发表时间:
2018
影响因子:
2.9
通讯作者:
Jie Han
Jie Han
中科院分区:
计算机科学3区
文献类型:
--
作者:
Siting Liu;Honglan Jiang;Leibo Liu;Jie Han

文献摘要

参考文献

被引文献

相似文献

梯度下降(GD)是机器学习中广泛使用的优化算法。本文提出了一种新的随机计算GD电路(SC-GDC),它将梯度信息编码到随机序列中。受神经元结构的启发,随机积分器被用于通过其“抑制”和“兴奋”输入来优化学习机中的权重。具体地,用于单极表示(或双极表示)的两个AND(或XNOR)门和一个随机积分器分别用于实现GD算法中的乘法和累加。因此,SC-GDC非常具有面积和功率效率。根据所提出的SC-GDC的公式,它在学习算法中提供了优化权重的无偏估计。然后,建议的SC-GDC用于实现最小均方算法和softmax回归。在类似的精度下,所提出的设计实现了超过<inline-formula><tex-math notation="LaTeX">30美元 与</tex-math></inline-formula>定点实现相比,每面积吞吐量(TPA)提高了10倍,每个训练样本消耗的能量不到13%。此外,一个有符号的SC-GDC提出了训练复杂的神经网络(NN)。结果表明,对于784-128-128-10全连接神经网络,带符号的SC-GDC产生与其固定点对应物相似的训练结果,同时实现了超过90%的能量节省和82%的训练时间减少超过<inline-formula><tex-math notation="LaTeX">50美元 TPA</tex-math></inline-formula>提高了10倍。
Gradient descent (GD) is a widely used optimization algorithm in machine learning. In this paper, a novel stochastic computing GD circuit (SC-GDC) is proposed by encoding the gradient information in stochastic sequences. Inspired by the structure of a neuron, a stochastic integrator is used to optimize the weights in a learning machine by its “inhibitory” and “excitatory” inputs. Specifically, two AND (or XNOR) gates for the unipolar representation (or the bipolar representation) and one stochastic integrator are, respectively, used to implement the multiplications and accumulations in a GD algorithm. Thus, the SC-GDC is very area- and power-efficient. As per the formulation of the proposed SC-GDC, it provides unbiased estimate of the optimized weights in a learning algorithm. The proposed SC-GDC is then used to implement a least-mean-square algorithm and a softmax regression. With a similar accuracy, the proposed design achieves more than <inline-formula> <tex-math notation="LaTeX">$30 imes $ </tex-math></inline-formula> improvement in throughput per area (TPA) and consumes less than 13% of the energy per training sample, compared with a fixed-point implementation. Moreover, a signed SC-GDC is proposed for training complex neural networks (NNs). It is shown that for a 784-128-128-10 fully connected NN, the signed SC-GDC produces a similar training result with its fixed-point counterpart, while achieving more than 90% energy saving and 82% reduction in training time with more than <inline-formula> <tex-math notation="LaTeX">$50 imes $ </tex-math></inline-formula> improvement in TPA.
DOI: 10.1109/tcsi.2018.2856513
发表时间: 2019-01-01
影响因子: 5.1
作者:
Jiang, Honglan;Liu, Leibo;Han, Jie
通讯作者: Han, Jie