Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model

Robust Knowledge Transfer via Hybrid Forward on the Teacher-Student Model
复制标题

DOI:
10.1609/aaai.v35i3.16358
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Liangchen Song;Jialian Wu;Ming Yang;Qian Zhang;Yuan Li;Junsong Yuan
Liangchen Song;Jialian Wu;Ming Yang;Qian Zhang;Yuan Li;Junsong Yuan
中科院分区:
其他
文献类型:
--
作者:
Liangchen Song;Jialian Wu;Ming Yang;Qian Zhang;Yuan Li;Junsong Yuan

文献摘要

被引文献

相似文献

当采用深度神经网络进行新的视觉任务时,通常的做法是从微调社区中一些现成的训练有素的网络模型开始。由于新任务可能需要使用新的域数据来训练不同的网络架构,因此利用现成的模型并不简单,通常需要大量的试错和参数调整。在本文中,我们将一个经过良好训练的模型表示为教师网络,将新任务的模型表示为学生网络。我们的目标是减轻将知识从教师网络转移到学生网络的工作,对他们的网络架构,域数据和任务定义之间的差距具有鲁棒性。具体来说,我们提出了一个混合的前向计划在训练教师-学生模型,交替更新层的学生模型的权重。我们的混合前向方案的关键优点是在训练中的知识转移损失和任务特定损失之间的动态平衡。我们证明了我们的方法对各种任务的有效性,例如,模型压缩,分割和检测,在各种知识转移设置。
When adopting deep neural networks for a new vision task, a common practice is to start with fine-tuning some off-the-shelf well-trained network models from the community. Since a new task may require training a different network architecture with new domain data, taking advantage of off-the-shelf models is not trivial and generally requires considerable try-and-error and parameter tuning. In this paper, we denote a well-trained model as a teacher network and a model for the new task as a student network. We aim to ease the efforts of transferring knowledge from the teacher to the student network, robust to the gaps between their network architectures, domain data, and task definitions. Specifically, we propose a hybrid forward scheme in training the teacher-student models, alternately updating layer weights of the student model. The key merit of our hybrid forward scheme is on the dynamical balance between the knowledge transfer loss and task specific loss in training. We demonstrate the effectiveness of our method on a variety of tasks, e.g., model compression, segmentation, and detection, under a variety of knowledge transfer settings.