L2L: A Highly Accurate Log_2_Lead Quantization of Pre-trained Neural Networks

L2L: A Highly Accurate Log_2_Lead Quantization of Pre-trained Neural Networks
复制标题

L2L:预训练神经网络的高精度 Log_2_Lead 量化

DOI:
10.23919/date48585.2020.9116373
复制
发表时间:
2020
期刊:
2020 Design, Automation & Test in Europe Conference & Exhibition (DATE)
影响因子:
--
通讯作者:
Akash Kumar
Akash Kumar
中科院分区:
--
文献类型:
--
作者:
Salim Ullah;Siddharth Gupta;K. Ahuja;Aruna Tiwari;Akash Kumar

文献摘要

被引文献

相似文献

深度神经网络是机器学习技术中的一种,在各种应用中得到越来越广泛的应用。然而,深度神经网络对内存和计算的巨大需求往往限制了其在嵌入式系统中的应用。最近的许多工作通过提出不同类型的数据量化方案来考虑这一问题。然而,这些技术中的大多数要么需要对深度神经网络进行量化后的再训练,要么在输出精度上有很大损失。本文提出了一种新的预训练深度神经网络参数量化方法。我们的技术显著地保持了参数的准确性,并且不需要对网络进行重新训练。与基于单精度浮点数的实现相比,对于使用ImageNet数据集的VGG16网络,我们提出的8位量化技术仅产生约1%的损失和约0.4%的TOP-1和TOP-5精度损失。
Deep Neural Networks are one of the machine learning techniques which are increasingly used in a variety of applications. However, the significantly high memory and computation demands of deep neural networks often limit their deployment on embedded systems. Many recent works have considered this problem by proposing different types of data quantization schemes. However, most of these techniques either require post-quantization retraining of deep neural networks or bear a significant loss in output accuracy. In this paper, we propose a novel quantization technique for parameters of pre-trained deep neural networks. Our technique significantly maintains the accuracy of the parameters and does not require retraining of the networks. Compared to the single-precision floating-point numbers-based implementation, our proposed 8-bit quantization technique generates only ~1% and the ~0.4%, loss in top-1 and top-5 accuracies respectively for VGG16 network using ImageNet dataset.