MetaDistiller: Network Self-Boosting via Meta-Learned Top-Down Distillation

MetaDistiller: Network Self-Boosting via Meta-Learned Top-Down Distillation
复制标题

DOI:
10.1007/978-3-030-58568-6_41
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Benlin Liu;Yongming Rao;Jiwen Lu;Jie Zhou;Cho-Jui Hsieh
Benlin Liu;Yongming Rao;Jiwen Lu;Jie Zhou;Cho-Jui Hsieh
中科院分区:
其他
文献类型:
--
作者:
Benlin Liu;Yongming Rao;Jiwen Lu;Jie Zhou;Cho-Jui Hsieh

文献摘要

被引文献

相似文献

知识蒸馏(KD)是学习紧凑模型的最流行的方法之一。然而,它仍然遭受高的时间和计算资源的顺序训练管道造成的要求。此外,由于兼容性的差距,来自较深模型的软目标通常不作为较浅模型的良好线索。在这项工作中,我们同时考虑这两个问题。具体来说,我们建议,更好的软目标,具有更高的兼容性,可以通过使用标签生成器融合的特征图从更深的阶段,在自上而下的方式,我们可以采用元学习技术来优化这个标签生成器。利用从模型的中间特征图中学习的软目标,我们可以实现更好的自增强的网络相比,国家的最先进的。实验进行了两个标准的分类基准,即CIFAR-100和ILSVRC 2012。我们测试各种网络架构,以显示我们的MetaDistiller的通用性。在两个数据集上的实验结果表明了该方法的有效性。
Knowledge Distillation (KD) has been one of the most popular methods to learn a compact model. However, it still suffers from high demand in time and computational resources caused by sequential training pipeline. Furthermore, the soft targets from deeper models do not often serve as good cues for the shallower models due to the gap of compatibility. In this work, we consider these two problems at the same time. Specifically, we propose that better soft targets with higher compatibility can be generated by using a label generator to fuse the feature maps from deeper stages in a top-down manner, and we can employ the meta-learning technique to optimize this label generator. Utilizing the soft targets learned from the intermediate feature maps of the model, we can achieve better self-boosting of the network in comparison with the state-of-the-art. The experiments are conducted on two standard classification benchmarks, namely CIFAR-100 and ILSVRC2012. We test various network architectures to show the generalizability of our MetaDistiller. The experiments results on two datasets strongly demonstrate the effectiveness of our method .