Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks

Divide and Conquer: Leveraging Intermediate Feature Representations for Quantized Training of Neural Networks
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Ahmed T. Elthakeb;Prannoy Pilligundla;H. Esmaeilzadeh
Ahmed T. Elthakeb;Prannoy Pilligundla;H. Esmaeilzadeh
中科院分区:
其他
文献类型:
--
作者:
Ahmed T. Elthakeb;Prannoy Pilligundla;H. Esmaeilzadeh

文献摘要

相似文献

现代神经网络的深层在输入通过网络传播时提取了相当丰富的特征集。本文旨在以最小的精度损失收获这些丰富的中间表示进行量化,同时显着降低DNN的内存占用和计算强度。本文通过师生范式(欣顿等人,2015)在一种新的设置中,利用DNN的特征提取能力进行更高精度的量化。因此,我们的算法在逻辑上将预训练的全精度DNN划分为多个部分,每个部分都暴露了中间特征,以便在量化域中独立训练一组学生。事实上,这种分而治之的策略使得每个学生部分的训练可以孤立进行,而所有这些独立训练的部分随后被缝合在一起以形成等效的完全量化的网络。我们的算法是一种面向知识蒸馏的分段方法,并且在一次知识蒸馏通过整个网络之前,不将中间表示视为预训练的提示(Romero等人,2015年)。在各种DNN(AlexNet、LeNet、MobileNet、ResNet-18、ResNet-20、SVHN和VGG-11)上的实验表明,这种称为DCQ(Divide and Conquer Quantization)的方法平均而言提高了最先进的量化训练技术DoReFa-Net(Zhou et al.,2016年),分别为21.6%和9.3%的二进制和三进制量化。此外,我们表明,将DCQ纳入现有的量化训练方法,导致提高精度相比,以前报道的多个国家的最先进的量化训练方法。
The deep layers of modern neural networks extract a rather rich set of features as an input propagates through the network. This paper sets out to harvest these rich intermediate representations for quantization with minimal accuracy loss while significantly reducing the memory footprint and compute intensity of the DNN. This paper utilizes knowledge distillation through teacher-student paradigm (Hinton et al., 2015) in a novel setting that exploits the feature extraction capability of DNNs for higher-accuracy quantization. As such, our algorithm logically divides a pretrained full-precision DNN to multiple sections, each of which exposes intermediate features to train a team of students independently in the quantized domain. This divide and conquer strategy, in fact, makes the training of each student section possible in isolation while all these independently trained sections are later stitched together to form the equivalent fully quantized network. Our algorithm is a sectional approach towards knowledge distillation and is not treating the intermediate representation as a hint for pretraining before one knowledge distillation pass over the entire network (Romero et al., 2015). Experiments on various DNNs (AlexNet, LeNet, MobileNet, ResNet-18, ResNet-20, SVHN and VGG-11) show that, this approach -- called DCQ (Divide and Conquer Quantization) -- on average, improves the performance of a state-of-the-art quantized training technique, DoReFa-Net (Zhou et al., 2016) by 21.6% and 9.3% for binary and ternary quantization, respectively. Additionally, we show that incorporating DCQ to existing quantized training methods leads to improved accuracies as compared to previously reported by multiple state-of-the-art quantized training methods.