Distilling from Similar Tasks for Transfer Learning on a Budget

Distilling from Similar Tasks for Transfer Learning on a Budget
复制标题

DOI:
10.1109/iccv51070.2023.01050
复制
发表时间:
2023-04
期刊:
2023 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Kenneth Borup;Cheng Perng Phoo;Bharath Hariharan
Kenneth Borup;Cheng Perng Phoo;Bharath Hariharan
中科院分区:
其他
文献类型:
--
作者:
Kenneth Borup;Cheng Perng Phoo;Bharath Hariharan

文献摘要

相似文献

我们解决了在有限的标签下获得高效而准确的识别系统的挑战。虽然识别模型随着模型大小和数据量的增加而改进,但计算机视觉的许多专业应用在训练和推理过程中都存在严重的资源限制。迁移学习是一种有效的解决方案,用于训练很少的标签,但通常以计算成本高昂的大基础模型微调为代价。我们建议通过从一组不同的源模型中进行半监督跨域蒸馏来减轻计算和准确性之间的这种不愉快的权衡。首先,我们展示了如何使用任务相似性度量来选择一个合适的源模型进行提取,并且一个好的选择过程对于目标模型的良好下游性能是必要的。我们称这种方法为DistillNearest。虽然有效,但DistillNearest假设单个源模型与目标任务匹配,但情况并非总是如此。为了缓解这一问题,我们提出了一种加权多源蒸馏方法,将在不同域上训练的多个源模型根据其与目标任务的相关性加权提取为单个有效模型(名为DistillWeighted)。我们的方法不需要访问源数据,只需要源模型的特征和伪标签。当目标是在计算约束下准确识别时,DistillNearest和DistillWeighted方法的性能都优于来自强ImageNet初始化的迁移学习以及最先进的半监督技术,如FixMatch。平均超过8个不同的目标任务,我们的多源方法分别比基线高出5.6%和4.5%。代码:github.com/Kennethborup/DistillWeighted
We address the challenge of getting efficient yet accurate recognition systems with limited labels. While recognition models improve with model size and amount of data, many specialized applications of computer vision have severe resource constraints both during training and inference. Transfer learning is an effective solution for training with few labels, however often at the expense of a computationally costly fine-tuning of large base models. We propose to mitigate this unpleasant trade-off between compute and accuracy via semi-supervised cross-domain distillation from a set of diverse source models. Initially, we show how to use task similarity metrics to select a single suitable source model to distill from, and that a good selection process is imperative for good downstream performance of a target model. We dub this approach DistillNearest. Though effective, DistillNearest assumes a single source model matches the target task, which is not always the case. To alleviate this, we propose a weighted multi-source distillation method to distill multiple source models trained on different domains weighted by their relevance for the target task into a single efficient model (named DistillWeighted). Our methods need no access to source data and merely need features and pseudo-labels of the source models. When the goal is accurate recognition under computational constraints, both DistillNearest and DistillWeighted approaches outperform both transfer learning from strong ImageNet initializations as well as state-of-the-art semisupervised techniques such as FixMatch. Averaged over 8 diverse target tasks our multi-source method outperforms the baselines by 5.6%-points and 4.5%-points, respectively. Code: github.com/Kennethborup/DistillWeighted