Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox Model

Neural Networks Are More Productive Teachers Than Human Raters: Active Mixup for Data-Efficient Knowledge Distillation From a Blackbox Model
复制标题

DOI:
10.1109/cvpr42600.2020.00157
复制
发表时间:
2020-03
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Dongdong Wang;Yandong Li;Liqiang Wang;Boqing Gong
Dongdong Wang;Yandong Li;Liqiang Wang;Boqing Gong
中科院分区:
其他
文献类型:
--
作者:
Dongdong Wang;Yandong Li;Liqiang Wang;Boqing Gong

文献摘要

被引文献

相似文献

我们研究如何通过以数据有效的方式从黑盒教师模型中提取知识来训练学生深度神经网络进行视觉识别。这个问题的进展可以显着减少学习高性能视觉识别模型对大规模数据集的依赖。有两个主要挑战。一是应尽量减少对教师模型的查询数量,以节省计算和/或财务成本。二是用于知识蒸馏的图像数量要少;否则,就违背了我们减少对大规模数据集依赖的期望。为了应对这些挑战,我们提出了一种融合混合和主动学习的方法。前者通过从原始图像的凸包采样的大量合成图像有效地增强了少数未标记图像,后者主动从池中为学生神经网络选择硬示例,并从教师模型中查询其标签。我们通过大量实验验证了我们的方法。
We study how to train a student deep neural network for visual recognition by distilling knowledge from a blackbox teacher model in a data-efficient manner. Progress on this problem can significantly reduce the dependence on large-scale datasets for learning high-performing visual recognition models. There are two major challenges. One is that the number of queries into the teacher model should be minimized to save computational and/or financial costs. The other is that the number of images used for the knowledge distillation should be small; otherwise, it violates our expectation of reducing the dependence on large-scale datasets. To tackle these challenges, we propose an approach that blends mixup and active learning. The former effectively augments the few unlabeled images by a big pool of synthetic images sampled from the convex hull of the original images, and the latter actively chooses from the pool hard examples for the student neural network and query their labels from the teacher model. We validate our approach with extensive experiments.