Feature Affinity Assisted Knowledge Distillation and Quantization of Deep Neural Networks on Label-Free Data

Feature Affinity Assisted Knowledge Distillation and Quantization of Deep Neural Networks on Label-Free Data
复制标题

DOI:
10.1109/access.2023.3297890
复制
发表时间:
2023-02
期刊:
影响因子:
3.9
通讯作者:
Zhijian Li;Biao Yang;Penghang Yin;Y. Qi;J. Xin
Zhijian Li;Biao Yang;Penghang Yin;Y. Qi;J. Xin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhijian Li;Biao Yang;Penghang Yin;Y. Qi;J. Xin

文献摘要

相似文献

本文提出了一种特征亲和力(FA)辅助的知识蒸馏(KD)方法来改进深度神经网络(DNN)的量化感知训练。DNN中间特征图上的FA损失起到了向学生传授解决方案的中间步骤的作用,而不是仅仅在传统的KD中给出最终答案,其中损失作用于输出级的网络逻辑。结合Logit损失和FA损失,我们通过在CIFAR-10/100和微小ImageNet数据集上的卷积网络实验发现,量化的学生网络比标记的地面真实数据受到更强的监督。由此产生的FA量化蒸馏(FAQD),经过200个历元的余弦退火调度器训练收敛,能够将无标签数据上的模型压缩到或超过其全精度对应模型的精度水平,这带来了直接的实际好处,因为预先训练的教师模型随时可用,并且未标记数据丰富。相比之下,数据标签通常既费力又昂贵。最后,我们提出并证明了一种快速特征亲和力损失函数的误差估计,它以较低的计算复杂度精确地逼近快速特征亲和力损失,这有助于提高高分辨率图像输入的训练速度。源代码可在以下网址获得:https://github.com/lzj994/FAQD
In this paper, we propose a feature affinity (FA) assisted knowledge distillation (KD) method to improve quantization-aware training of deep neural networks (DNN). The FA loss on intermediate feature maps of DNNs plays the role of teaching middle steps of a solution to a student instead of only giving final answers in the conventional KD where the loss acts on the network logits at the output level. Combining logit loss and FA loss, we found via convolutional network experiments on CIFAR-10/100, and Tiny ImageNet data sets that the quantized student network receives stronger supervision than from the labeled ground-truth data. The resulting FA quantization-distillation (FAQD), trained to convergence with a cosine annealing scheduler for 200 epochs, is capable of compressing models on label-free data up to or exceeding the accuracy levels of their full precision counterparts, which brings immediate practical benefits as pre-trained teacher models are readily available and unlabeled data are abundant. In contrast, data labeling is often laborious and expensive. Finally, we propose and prove error estimates for a fast feature affinity (FFA) loss function that accurately approximates FA loss at a lower order of computational complexity, which helps speed up training for high resolution image input. Source codes are available at: https://github.com/lzj994/FAQD