Gradient Descent Using Stochastic Circuits for Efficient Training of Learning Machines
Gradient Descent Using Stochastic Circuits for Efficient Training of Learning Machines
复制标题
使用随机电路的梯度下降来有效训练学习机
DOI:
--
复制
发表时间:
2018
影响因子:
2.9
通讯作者:
Jie Han
中科院分区:
文献类型:
--
作者:
Siting Liu;Honglan Jiang;Leibo Liu;Jie Han
Gradient descent (GD) is a widely used optimization algorithm in machine learning. In this paper, a novel stochastic computing GD circuit (SC-GDC) is proposed by encoding the gradient information in stochastic sequences. Inspired by the structure of a neuron, a stochastic integrator is used to optimize the weights in a learning machine by its “inhibitory” and “excitatory” inputs. Specifically, two AND (or XNOR) gates for the unipolar representation (or the bipolar representation) and one stochastic integrator are, respectively, used to implement the multiplications and accumulations in a GD algorithm. Thus, the SC-GDC is very area- and power-efficient. As per the formulation of the proposed SC-GDC, it provides unbiased estimate of the optimized weights in a learning algorithm. The proposed SC-GDC is then used to implement a least-mean-square algorithm and a softmax regression. With a similar accuracy, the proposed design achieves more than <inline-formula> <tex-math notation="LaTeX">$30 imes $ </tex-math></inline-formula> improvement in throughput per area (TPA) and consumes less than 13% of the energy per training sample, compared with a fixed-point implementation. Moreover, a signed SC-GDC is proposed for training complex neural networks (NNs). It is shown that for a 784-128-128-10 fully connected NN, the signed SC-GDC produces a similar training result with its fixed-point counterpart, while achieving more than 90% energy saving and 82% reduction in training time with more than <inline-formula> <tex-math notation="LaTeX">$50 imes $ </tex-math></inline-formula> improvement in TPA.
DOI:
10.1109/tcsi.2018.2856513
发表时间:
2019-01-01
影响因子:
5.1
作者:
Jiang, Honglan;Liu, Leibo;Han, Jie
通讯作者:
Han, Jie