BERT Learns to Teach: Knowledge Distillation with Meta Learning

BERT Learns to Teach: Knowledge Distillation with Meta Learning
复制标题

DOI:
10.18653/v1/2022.acl-long.485
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Wangchunshu Zhou;Canwen Xu;Julian McAuley
Wangchunshu Zhou;Canwen Xu;Julian McAuley
中科院分区:
其他
文献类型:
--
作者:
Wangchunshu Zhou;Canwen Xu;Julian McAuley

文献摘要

被引文献

相似文献

我们提出了基于元学习的知识蒸馏(MetaDistil),这是一种简单但有效的替代传统知识蒸馏(KD)方法的方法,其中教师模型在培训期间是固定的。我们发现,在元学习框架下,教师网络可以通过提取学生网络的性能反馈,更好地向学生网络传递知识(即学会教学)。此外,我们还引入了一种先导更新机制来改善元学习算法中内学习者和元学习者之间的一致性,该算法关注于改进的内学习者。在不同的基准测试上的实验表明,与传统的KD算法相比,MetaDistil算法有明显的改进,并且对不同的学生容量和超参数的选择不那么敏感,便于在不同的任务和模型上使用KD。
We present Knowledge Distillation with Meta Learning (MetaDistil), a simple yet effective alternative to traditional knowledge distillation (KD) methods where the teacher model is fixed during training. We show the teacher network can learn to better transfer knowledge to the student network (i.e., learning to teach) with the feedback from the performance of the distilled student network in a meta learning framework. Moreover, we introduce a pilot update mechanism to improve the alignment between the inner-learner and meta-learner in meta learning algorithms that focus on an improved inner-learner. Experiments on various benchmarks show that MetaDistil can yield significant improvements compared with traditional KD algorithms and is less sensitive to the choice of different student capacity and hyperparameters, facilitating the use of KD on different tasks and models.