Dual-Level Knowledge Distillation via Knowledge Alignment and Correlation

Dual-Level Knowledge Distillation via Knowledge Alignment and Correlation
复制标题

基于知识对齐和关联的双层知识蒸馏

DOI:
10.1109/tnnls.2022.3190166
复制
发表时间:
2022-07
影响因子:
10.4
通讯作者:
Fei Ding;Yin Yang;Hongxin Hu;V. Krovi;Feng Luo
Fei Ding;Yin Yang;Hongxin Hu;V. Krovi;Feng Luo
中科院分区:
计算机科学1区
文献类型:
--
作者:
Fei Ding;Yin Yang;Hongxin Hu;V. Krovi;Feng Luo

文献摘要

相似文献

知识蒸馏已成为一种广泛使用的模型压缩和知识传递技术。我们发现,标准的KD方法通过类原型间接地对单个样本进行知识比对,而忽略了不同样本之间的结构性知识,即知识相关性。虽然目前基于对比学习的蒸馏方法可以分解为知识对齐和关联,但它们的关联目标不希望将来自同一类的样本的表示分开,导致蒸馏结果较差。为了提高精馏性能,本文提出了一种新的知识关联目标,并引入了双层知识精馏(DLKD),它明确地将知识对齐和关联结合在一起,而不是使用单一的对比目标。我们证明了知识对齐和关联对于提高精馏性能是必要的。特别是,知识关联可以作为学习广义表示的一种有效的正则化。所提出的DLKD是任务不可知和模型不可知的,并且能够从受监督或自我监督的预训教师向学生有效地转移知识。实验表明,DLKD在大量实验环境中的性能优于其他最先进的方法,包括:1)预训练策略;2)网络体系结构;3)数据集;4)任务。
Knowledge distillation (KD) has become a widely used technique for model compression and knowledge transfer. We find that the standard KD method performs the knowledge alignment on an individual sample indirectly via class prototypes and neglects the structural knowledge between different samples, namely, knowledge correlation. Although recent contrastive learning-based distillation methods can be decomposed into knowledge alignment and correlation, their correlation objectives undesirably push apart representations of samples from the same class, leading to inferior distillation results. To improve the distillation performance, in this work, we propose a novel knowledge correlation objective and introduce the dual-level knowledge distillation (DLKD), which explicitly combines knowledge alignment and correlation together instead of using one single contrastive objective. We show that both knowledge alignment and correlation are necessary to improve the distillation performance. In particular, knowledge correlation can serve as an effective regularization to learn generalized representations. The proposed DLKD is task-agnostic and model-agnostic, and enables effective knowledge transfer from supervised or self-supervised pretrained teachers to students. Experiments show that DLKD outperforms other state-of-the-art methods on a large number of experimental settings including: 1) pretraining strategies; 2) network architectures; 3) datasets; and 4) tasks.