Scaffolding a Student to Instill Knowledge

Scaffolding a Student to Instill Knowledge
复制标题

DOI:
--
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Instill Knowledge;Anil Kag;D. A. Acar;Aditya Gangrade;Venkatesh Saligrama
Instill Knowledge;Anil Kag;D. A. Acar;Aditya Gangrade;Venkatesh Saligrama
中科院分区:
其他
文献类型:
--
作者:
Instill Knowledge;Anil Kag;D. A. Acar;Aditya Gangrade;Venkatesh Saligrama

文献摘要

相似文献

我们提出了一种新的知识蒸馏(KD)方法,在学生的能力明显小于教师的情况下,有选择地将教师知识灌输到学生模型中。在香草式KD中,教师主要是为学生设定一个预测性目标,让学生遵循,我们认为由于学生缺乏能力,这个目标过于乐观。我们开发了一种新的脚手架方案,其中教师除了设置预测目标外,还通过审查难以学习的示例来支撑学生的预测。学生模型使用与老师的软最大预测相同的信息作为输入,从这个意义上说,我们的建议可以被视为香草KD的自然变体。我们在合成示例中表明,审查硬示例会平滑学生的损失情况,以便学生遇到更少的局部最小值。因此,它具有良好的泛化性质。对于香草KD,我们实现了改进的性能,并且可以与在基准数据集上利用特征匹配的更具侵入性的技术相媲美。
We propose a novel knowledge distillation (KD) method to selectively instill teacher knowledge into a student model motivated by situations where the student’s capacity is significantly smaller than that of the teachers. In vanilla KD, the teacher primarily sets a predictive target for the student to follow, and we posit that this target is overly optimistic due to the student’s lack of capacity. We develop a novel scaffolding scheme where the teacher, in addition to setting a predictive target, also scaffolds the student’s prediction by censoring hard-to-learn examples. The student model utilizes the same information as the teacher’s soft-max predictions as inputs, and in this sense, our proposal can be viewed as a natural variant of vanilla KD. We show on synthetic examples that censoring hard-examples leads to smoothening the student’s loss landscape so that the student encounters fewer local minima. As a result, it has good generalization properties. Against vanilla KD, we achieve improved performance and are comparable to more intrusive techniques that leverage feature matching on benchmark datasets.