Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation

Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation
复制标题

比较知识蒸馏中的 Kullback-Leibler 散度和均方误差损失

DOI:
10.24963/ijcai.2021/362
复制
发表时间:
2021
期刊:
Canadian Journal of Mathematics
影响因子:
--
通讯作者:
Se
Se
中科院分区:
--
文献类型:
--
作者:
Taehyeon Kim;Jaehoon Oh;Nakyil Kim;Sangwook Cho;Se

文献摘要

被引文献

相似文献

知识蒸馏(KD),将知识从一个笨重的教师模型转移到一个轻量级的学生模型,已被研究设计有效的神经结构。通常,KD的目标函数是教师模型和学生模型的软化概率分布与温度标度超参数τ之间的Kullback-Leibler(KL)发散损失。尽管它被广泛使用,但很少有研究讨论这种软化如何影响概括。在这里,我们从理论上表明,当τ增加时,KL发散损失集中在logit匹配上,当τ变为0时,KL发散损失集中在标签匹配上,并且经验上表明,logit匹配与性能改善一般呈正相关。从这个观察中,我们考虑一个直观的KD损失函数,logit向量之间的均方误差(MSE),这样学生模型就可以直接学习教师模型的logit。MSE损失优于KL发散损失,这可以通过两种损失之间的倒数第二层表示差异来解释。此外,我们表明,顺序蒸馏可以提高性能和KD,特别是使用小τ的KL发散损失,减轻标签噪声。复制实验的代码可在https://github.com/jhoon-oh/kd_data/在线公开获得。
Knowledge distillation (KD), transferring knowledge from a cumbersome teacher model to a lightweight student model, has been investigated to design efficient neural architectures. Generally, the objective function of KD is the Kullback-Leibler (KL) divergence loss between the softened probability distributions of the teacher model and the student model with the temperature scaling hyperparameter τ. Despite its widespread use, few studies have discussed how such softening influences generalization. Here, we theoretically show that the KL divergence loss focuses on the logit matching when τ increases and the label matching when τ goes to 0 and empirically show that the logit matching is positively correlated to performance improvement in general. From this observation, we consider an intuitive KD loss function, the mean squared error (MSE) between the logit vectors, so that the student model can directly learn the logit of the teacher model. The MSE loss outperforms the KL divergence loss, explained by the penultimate layer representations difference between the two losses. Furthermore, we show that sequential distillation can improve performance and that KD, using the KL divergence loss with small τ particularly, mitigates the label noise. The code to reproduce the experiments is publicly available online at https://github.com/jhoon-oh/kd_data/.