Trained Ternary Quantization

Trained Ternary Quantization
复制标题

DOI:
--
复制
发表时间:
2016-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally
中科院分区:
其他
文献类型:
--
作者:
Chenzhuo Zhu;Song Han;Huizi Mao;W. Dally

文献摘要

被引文献

相似文献

深度神经网络广泛用于机器学习应用。然而,大型神经网络模型的部署可能难以部署在具有有限功率预算的移动的设备上。为了解决这个问题,我们提出了训练的三进制量化(TTQ),一种可以将神经网络中的权重精度降低到三进制值的方法。这种方法的精度下降非常小,甚至可以提高CIFAR-10和ImageNet上AlexNet的一些模型(32层,44层,56层ResNet)的精度。我们的AlexNet模型是从头开始训练的,这意味着它和训练正常的全精度模型一样容易。我们强调了我们训练的量化方法,它可以学习三进制值和三进制赋值。在推理过程中,只需要三进制值(2位权重)和缩放因子,因此我们的模型比全精度模型小近16倍。我们的三进制模型也可以被视为稀疏的二进制权重网络,可以使用自定义电路来加速。在CIFAR-10上的实验表明,训练量化方法得到的三值模型比ResNet-32、44、56的全精度模型分别高出0.04%、0.16%、0.36%。在ImageNet上,我们的模型比全精度AlexNet模型高出0.3%的Top-1精度,比之前的三元模型高出3%。
Deep neural networks are widely used in machine learning applications. However, the deployment of large neural networks models can be difficult to deploy on mobile devices with limited power budgets. To solve this problem, we propose Trained Ternary Quantization (TTQ), a method that can reduce the precision of weights in neural networks to ternary values. This method has very little accuracy degradation and can even improve the accuracy of some models (32, 44, 56-layer ResNet) on CIFAR-10 and AlexNet on ImageNet. And our AlexNet model is trained from scratch, which means it's as easy as to train normal full precision model. We highlight our trained quantization method that can learn both ternary values and ternary assignment. During inference, only ternary values (2-bit weights) and scaling factors are needed, therefore our models are nearly 16x smaller than full-precision models. Our ternary models can also be viewed as sparse binary weight networks, which can potentially be accelerated with custom circuit. Experiments on CIFAR-10 show that the ternary models obtained by trained quantization method outperform full-precision models of ResNet-32,44,56 by 0.04%, 0.16%, 0.36%, respectively. On ImageNet, our model outperforms full-precision AlexNet model by 0.3% of Top-1 accuracy and outperforms previous ternary models by 3%.