Training generalizable quantized deep neural nets

Training generalizable quantized deep neural nets
复制标题

DOI:
10.1016/j.eswa.2022.118736
复制
发表时间:
2022-10
期刊:
Expert Syst. Appl.
影响因子:
--
通讯作者:
Charles Hernandez;Bijan Taslimi;H. Lee;Hongcheng Liu;P. Pardalos
Charles Hernandez;Bijan Taslimi;H. Lee;Hongcheng Liu;P. Pardalos
中科院分区:
其他
文献类型:
--
作者:
Charles Hernandez;Bijan Taslimi;H. Lee;Hongcheng Liu;P. Pardalos

文献摘要

相似文献

虽然文献中已经提出了许多训练量化深度学习模型的实用方法,但这些方法的理论推广结果存在着严重的差距。尽管经验证据通常表明深度学习架构对训练过程的变化具有较高的容忍度,但现有的理论泛化分析通常取决于训练算法的具体设计,例如随机梯度下降(SGD)。这种特殊性使得这种普遍性结果不适用于量化深度学习模型的情况。鉴于这种临界真空,本文提供了几种几乎与算法无关的结果,以确保量化神经网络在不同最优水平下的通用性。这些结果包括确保泛化性能的可计算的量化局部解的特征,以及可证明收敛于此类局部解的算法。
While a number of practical methods for training quantized DL models have been presented in the literature, there exists a critical gap in the theoretical generalizability results for such approaches. Although empirical evidence often suggests a high tolerance of DL architectures to variations of training procedures, existing theoretical generalization analyses are often contingent on the specific designs of training algorithms, e.g., in stochastic gradient descent (SGD). This specialization makes such generalizability results inapplicable to the case of quantized DL models. In view of this critical vacuum, this paper provides several almost-algorithm-independent results to ensure the generalizability of a quantized neural network at different levels of optimality. These results include the characterizations of a computable, quantized local solution that ensures the generalization performance and an algorithm that is provably convergent to such a local solution.