Adaptive Quantization for Deep Neural Network

Adaptive Quantization for Deep Neural Network
复制标题

DOI:
10.1609/aaai.v32i1.11623
复制
发表时间:
2017-12
期刊:
J. Inf. Sci. Eng.
影响因子:
--
通讯作者:
Yiren Zhou;Seyed-Mohsen Moosavi-Dezfooli;Ngai-Man Cheung;P. Frossard
Yiren Zhou;Seyed-Mohsen Moosavi-Dezfooli;Ngai-Man Cheung;P. Frossard
中科院分区:
其他
文献类型:
--
作者:
Yiren Zhou;Seyed-Mohsen Moosavi-Dezfooli;Ngai-Man Cheung;P. Frossard

文献摘要

被引文献

相似文献

近年来,深度神经网络(DNN)在各种应用中得到了迅速发展,其结构也日益复杂。这些DNN的性能提升通常伴随着高计算成本和大内存消耗,这对于移动平台来说可能是负担不起的。深度模型量化可以用于降低DNN的计算和存储开销,以及在移动设备上部署复杂的DNN。在这项工作中,我们提出了一个深层模型量化的优化框架。首先,我们提出了一种度量来估计各个层的参数量化误差对整体模型预测精度的影响。然后,我们提出了一种基于该度量的优化过程来为每一层寻找最优的量化位宽。这是首次从理论上分析各层参数量化误差与模型精度之间的关系。新的量化算法的性能优于以往的量化优化方法,在相同的模型预测精度下,与等比特宽度的量化算法相比,压缩比提高了20-40%。
In recent years Deep Neural Networks (DNNs) have been rapidly developed in various applications, together with increasingly complex architectures. The performance gain of these DNNs generally comes with high computational costs and large memory consumption, which may not be affordable for mobile platforms. Deep model quantization can be used for reducing the computation and memory costs of DNNs, and deploying complex DNNs on mobile equipment. In this work, we propose an optimization framework for deep model quantization. First, we propose a measurement to estimate the effect of parameter quantization errors in individual layers on the overall model prediction accuracy. Then, we propose an optimization process based on this measurement for finding optimal quantization bit-width for each layer. This is the first work that theoretically analyse the relationship between parameter quantization errors of individual layers and model accuracy. Our new quantization algorithm outperforms previous quantization optimization methods, and achieves 20-40% higher compression rate compared to equal bit-width quantization at the same model prediction accuracy.