Post-training Quantization for Neural Networks with Provable Guarantees

Post-training Quantization for Neural Networks with Provable Guarantees
复制标题

DOI:
10.1137/22m1511709
复制
发表时间:
2022-01
期刊:
SIAM J. Math. Data Sci.
影响因子:
--
通讯作者:
Jinjie Zhang;Yixuan Zhou;Rayan Saab
Jinjie Zhang;Yixuan Zhou;Rayan Saab
中科院分区:
其他
文献类型:
--
作者:
Jinjie Zhang;Yixuan Zhou;Rayan Saab

文献摘要

被引文献

相似文献

虽然神经网络在广泛的应用中取得了显著的成功,但在资源受限的硬件中实现它们仍然是一个深入研究的领域。通过用量化的(例如,4位或二进制)对应,实现了计算成本、存储器和功耗的大量节省。为此,我们推广了一种基于贪婪路径跟踪机制的训练后神经网络量化方法GPFQ。除其他外,我们提出了修改,以促进稀疏的权重,并严格分析相关的错误。此外,我们的误差分析扩展了先前GPFQ工作的结果,以处理一般的量化字母表,表明对于量化单层网络,相对平方误差基本上在权重数量上线性衰减-即,过度参数化。我们的结果适用于一系列输入分布,并且适用于全连接和卷积架构,从而也扩展了以前的结果。为了对该方法进行经验评估,我们使用几种常见的架构,每种架构的权重都很少,并在ImageNet上进行测试,与未量化的模型相比,准确性只有轻微的损失。我们还表明,标准的修改,如偏差校正和混合精度量化,进一步提高精度。
While neural networks have been remarkably successful in a wide array of applications, implementing them in resource-constrained hardware remains an area of intense research. By replacing the weights of a neural network with quantized (e.g., 4-bit, or binary) counterparts, massive savings in computation cost, memory, and power consumption are attained. To that end, we generalize a post-training neural-network quantization method, GPFQ, that is based on a greedy path-following mechanism. Among other things, we propose modifications to promote sparsity of the weights, and rigorously analyze the associated error. Additionally, our error analysis expands the results of previous work on GPFQ to handle general quantization alphabets, showing that for quantizing a single-layer network, the relative square error essentially decays linearly in the number of weights -- i.e., level of over-parametrization. Our result holds across a range of input distributions and for both fully-connected and convolutional architectures thereby also extending previous results. To empirically evaluate the method, we quantize several common architectures with few bits per weight, and test them on ImageNet, showing only minor loss of accuracy compared to unquantized models. We also demonstrate that standard modifications, such as bias correction and mixed precision quantization, further improve accuracy.