Blended coarse gradient descent for full quantization of deep neural networks

Blended coarse gradient descent for full quantization of deep neural networks
复制标题

DOI:
10.1007/s40687-018-0177-6
复制
发表时间:
2018-08
影响因子:
1.2
通讯作者:
Penghang Yin;Shuai Zhang;J. Lyu;S. Osher;Y. Qi;J. Xin
Penghang Yin;Shuai Zhang;J. Lyu;S. Osher;Y. Qi;J. Xin
中科院分区:
数学3区
文献类型:
--
作者:
Penghang Yin;Shuai Zhang;J. Lyu;S. Osher;Y. Qi;J. Xin

文献摘要

相似文献

与常规的全精度神经网络相比,量化深度神经网络(QDNN)具有更低的存储容量和更快的推理速度,因此具有很大的吸引力。为了保持相同的性能水平,特别是在低位宽时,必须对QDNN进行重新训练。他们的训练涉及分段恒定激活函数和离散权重;因此,出现了数学挑战。我们引入了粗梯度的概念,提出了混合粗梯度下降(BCGD)算法,用于训练完全量化的神经网络。粗梯度一般不是任何函数的梯度,而是一种人工上升方向。BCGD的权重更新是通过对全精度权重的加权平均值及其量化(所谓的混合)进行粗梯度校正来进行的,这产生了目标值的充分下降,从而加速了训练。我们的实验表明,这种简单的混合技术对于二值化等极低比特宽度的量化非常有效。在ResNet-18 for ImageNet分类任务的完全量化中,BCGD在所有层的二进制权重和4位自适应激活的情况下提供了64.36%的TOP-1准确率。如果第一层和最后一层中的权重保持完全精度,则该数字将增加到65.46%。作为理论证明,我们给出了具有高斯输入数据的两层线性神经网络模型的粗梯度下降的收敛分析,并证明了期望的粗梯度与潜在的真实梯度正相关。
Quantized deep neural networks (QDNNs) are attractive due to their much lower memory storage and faster inference speed than their regular full-precision counterparts. To maintain the same performance level especially at low bit-widths, QDNNs must be retrained. Their training involves piecewise constant activation functions and discrete weights; hence, mathematical challenges arise. We introduce the notion of coarse gradient and propose the blended coarse gradient descent (BCGD) algorithm, for training fully quantized neural networks. Coarse gradient is generally not a gradient of any function but an artificial ascent direction. The weight update of BCGD goes by coarse gradient correction of a weighted average of the full-precision weights and their quantization (the so-called blending), which yields sufficient descent in the objective value and thus accelerates the training. Our experiments demonstrate that this simple blending technique is very effective for quantization at extremely low bit-width such as binarization. In full quantization of ResNet-18 for ImageNet classification task, BCGD gives 64.36% top-1 accuracy with binary weights across all layers and 4-bit adaptive activation. If the weights in the first and last layers are kept in full precision, this number increases to 65.46%. As theoretical justification, we show convergence analysis of coarse gradient descent for a two-linear-layer neural network model with Gaussian input data and prove that the expected coarse gradient correlates positively with the underlying true gradient.