KTAN: Knowledge Transfer Adversarial Network

KTAN: Knowledge Transfer Adversarial Network
复制标题

DOI:
10.1109/ijcnn48605.2020.9207235
复制
发表时间:
2018-10
期刊:
2020 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Peiye Liu;Wu Liu;Huadong Ma;Tao Mei;Mingoo Seok
Peiye Liu;Wu Liu;Huadong Ma;Tao Mei;Mingoo Seok
中科院分区:
其他
文献类型:
--
作者:
Peiye Liu;Wu Liu;Huadong Ma;Tao Mei;Mingoo Seok

文献摘要

被引文献

相似文献

首创知识蒸馏,将大型教师深度网络的泛化能力转移到轻量级的学生网络。学生网络可以保留教师网络的高质量,同时表现出较低的计算复杂度和存储要求,这对于在资源受限的移动设备上部署深度卷积神经网络很有吸引力。然而,大多数现有方法都专注于迁移教师网络中 softmax 层的概率分布,而忽略了中间表示。然而,我们发现,与仅概率分布相比,中间表示对于学生网络更好地理解转移泛化至关重要。因此,在本文中,我们提出了一种知识转移对抗网络方法,该方法全面考虑教师网络的中间表示和概率分布。为了迁移中间表示的知识,我们将高级教师特征图设置为目标,该方法将朝着该目标训练学生特征图。此外,为了支持学生网络的各种结构,我们安排了一个新颖的教师到学生层。最后,所提出的方法采用对抗性学习过程。具体来说,它包括一个鉴别器网络,以在学生网络的训练过程中充分利用特征图的空间相关性。实验结果表明,所提出的方法可以显着提高学生网络在图像分类和目标检测这两个重要视觉任务上的性能。
Knowledge distillation was pioneered to transfer the generalization ability of a large teacher deep network to a light-weight student network. The student network can retain the high quality of the teacher network, yet exhibiting low computational complexity and storage requirement, which is attractive for deploying a deep convolution neural network on a resource-constrained mobile device. However, most of the existing methods focus on transferring the probability distribution of a softmax layer in a teacher network and neglect the intermediate representations. However, we find that the intermediate representation is critical for a student network to better understand the transferred generalization as compared to the probability distribution only. In this paper, therefore, we propose such a knowledge transfer adversarial network method which holistically considers both intermediate representations and probability distributions of a teacher network. To transfer the knowledge of intermediate representations, we set high-level teacher feature maps as a target, toward which the method trains student feature maps. Furthermore, to support various structures of a student network, we arrange a novel teacher-to-student layer. Finally, the proposed method employs an adversarial learning process. Specifically, it includes a discriminator network to fully exploit the spatial correlation of feature maps during the training process of a student network. The experimental results demonstrate that the proposed method can significantly improve the performance of a student network on two important vision tasks, image classification and object detection.