Online Knowledge Distillation by Temporal-Spatial Boosting

Online Knowledge Distillation by Temporal-Spatial Boosting
复制标题

DOI:
10.1109/wacv51458.2022.00354
复制
发表时间:
2022-01
期刊:
2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Chengcheng Li;Zi Wang;Hairong Qi
Chengcheng Li;Zi Wang;Hairong Qi
中科院分区:
其他
文献类型:
--
作者:
Chengcheng Li;Zi Wang;Hairong Qi

文献摘要

相似文献

在线知识蒸馏(KD)以对等教学的方式从头开始相互训练一组学生网络,无需预先训练的教师模型。然而,来自同行的监督可能会很嘈杂,特别是在培训的早期阶段。在本文中,我们提出了一种新的方法,在线知识蒸馏的时空推进(TSB)。该方法通过时间累加器和空间积分器两个模块构建上级“教师”。具体而言,时间累加器在训练期间利用网络的先前输出,并在所有类别上产生代表性预测。我们不是像vanilla online KD那样仅仅模仿其他网络的输出,而是进一步提出了所谓的空间集成器,它整合了所有网络学习到的知识,并产生了更强大的指导者。这两个模块的操作简单明了,可以在训练过程中快速有效地计算。该方法可以提高有效知识的传递效率,同时稳定训练过程。在各种基准数据集和网络结构上的实验结果验证了该方法的有效性。
Online knowledge distillation (KD) mutually trains a group of student networks from scratch in a peer-teaching manner, eliminating the need for pre-trained teacher models. However, supervision from peers can be noisy, especially in the early stage of training. In this paper, we propose a novel method for online knowledge distillation by temporal-spatial boosting (TSB). The proposed method constructs superior "teachers" with two modules, temporal accumulator and spatial integrator. Specifically, the temporal accumulator leverages the previous outputs of networks during training and produces a representative prediction over all classes. Instead of merely imitating the outputs of other networks as in vanilla online KD, we further propose the so-called spatial integrator that consolidates the knowledge learned by all networks and yields a stronger instructor. The operations of these two modules are simple and straightforward, which can be computed efficiently on the fly during training. The proposed method can improve the efficiency of transferring effective knowledge as well as stabilize the training process. Experimental results on various benchmark datasets and network structures validate the effectiveness of the proposed method over the state-of-the-art.