Recurrence of optimum for training weight and activation quantized networks

Recurrence of optimum for training weight and activation quantized networks
复制标题

训练权重和激活量化网络的最佳重现

DOI:
10.1016/j.acha.2022.07.006
复制
发表时间:
2023
影响因子:
2.5
通讯作者:
Xin, Jack
Xin, Jack
中科院分区:
数学1区
文献类型:
--
作者:
Long, Ziang;Yin, Penghang;Xin, Jack

文献摘要

相似文献

深度神经网络(DNN)被量化,以在资源受限的平台上进行有效的推理。然而,训练具有低精度权重和激活的深度学习模型涉及到一项苛刻的优化任务,这需要最小化受离散集约束的阶段损失函数。虽然已经提出了许多训练方法,但现有的DNN全量化研究大多是经验性的。从理论的角度来看,我们研究的实用技术,克服网络量化的组合性质。具体来说,我们研究了一种简单而强大的投影梯度类算法,用于量化双层卷积网络,通过在量化权重处评估的损失函数的伪梯度(所谓的粗梯度)的负方向上重复移动浮点权重处的一步。我们首次证明了在温和的条件下,量化的权重序列递归地访问用于训练完全量化网络的离散最小化问题的全局最优解。我们还展示了训练量化深度网络中权重演化的递归现象的数值证据。
Deep neural networks (DNNs) are quantized for efficient inference on resource-constrained platforms. However, training deep learning models with low-precision weights and activations involves a demanding optimization task, which calls for minimizing a stage-wise loss function subject to a discrete set-constraint. While numerous training methods have been proposed, existing studies for full quantization of DNNs are mostly empirical. From a theoretical point of view, we study practical techniques for overcoming the combinatorial nature of network quantization. Specifically, we investigate a simple yet powerful projected gradient-like algorithm for quantizing two-layer convolutional networks, by repeatedly moving one step at float weights in the negative direction of a heuristicfakegradient of the loss function (so-called coarse gradient) evaluated at quantized weights. For the first time, we prove that under mild conditions, the sequence of quantized weights recurrently visit the global optimum of the discrete minimization problem for training a fully quantized network. We also show numerical evidence of the recurrence phenomenon of weight evolution in training quantized deep networks.