Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective

Rethinking Soft Labels for Knowledge Distillation: A Bias-Variance Tradeoff Perspective
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Helong Zhou;Liangchen Song;Jiajie Chen;Ye Zhou;Guoli Wang;Junsong Yuan;Qian Zhang
Helong Zhou;Liangchen Song;Jiajie Chen;Ye Zhou;Guoli Wang;Junsong Yuan;Qian Zhang
中科院分区:
其他
文献类型:
--
作者:
Helong Zhou;Liangchen Song;Jiajie Chen;Ye Zhou;Guoli Wang;Junsong Yuan;Qian Zhang

文献摘要

相似文献

知识蒸馏是一种有效的方法,利用一个训练有素的网络或他们的集合,被称为教师,来指导学生网络的训练。教师网络的输出被用作软标签,用于监督新网络的训练。最近的研究\citep{muller2019does,yuan2020revisiting}揭示了软标签的一个有趣的特性,使标签变软可以很好地正则化学生网络。从统计学习的角度来看,正则化的目的是减少方差,但是对于带有软标签的训练,偏差和方差是如何变化的并不清楚。在本文中,我们研究了带有软标签的蒸馏带来的偏差-方差权衡。具体来说,我们观察到,在训练过程中,偏差-方差权衡随样本而变化。此外,在相同的蒸馏温度设置下,我们观察到蒸馏性能与某些特定样本的数量呈负相关,这些样本被称为正则化样本,因为这些样本导致偏差增加而方差减小。然而,我们的经验发现完全滤除正则化样本也会使蒸馏性能恶化。我们的发现启发我们提出了新的加权软标签,以帮助网络自适应地处理样本的偏差-方差权衡。在标准评价基准上的实验验证了该方法的有效性。我们的代码可在\url{https://github.com/bellymonster/Weighted-Soft-Label-Distillation}上获得。
Knowledge distillation is an effective approach to leverage a well-trained network or an ensemble of them, named as the teacher, to guide the training of a student network. The outputs from the teacher network are used as soft labels for supervising the training of a new network. Recent studies \citep{muller2019does,yuan2020revisiting} revealed an intriguing property of the soft labels that making labels soft serves as a good regularization to the student network. From the perspective of statistical learning, regularization aims to reduce the variance, however how bias and variance change is not clear for training with soft labels. In this paper, we investigate the bias-variance tradeoff brought by distillation with soft labels. Specifically, we observe that during training the bias-variance tradeoff varies sample-wisely. Further, under the same distillation temperature setting, we observe that the distillation performance is negatively associated with the number of some specific samples, which are named as regularization samples since these samples lead to bias increasing and variance decreasing. Nevertheless, we empirically find that completely filtering out regularization samples also deteriorates distillation performance. Our discoveries inspired us to propose the novel weighted soft labels to help the network adaptively handle the sample-wise bias-variance tradeoff. Experiments on standard evaluation benchmarks validate the effectiveness of our method. Our code is available at \url{https://github.com/bellymonster/Weighted-Soft-Label-Distillation}.