LiT: Zero-Shot Transfer with Locked-image text Tuning

LiT: Zero-Shot Transfer with Locked-image text Tuning
复制标题

DOI:
10.1109/cvpr52688.2022.01759
复制
发表时间:
2021-11
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Xiaohua Zhai;Xiao Wang;Basil Mustafa;A. Steiner;Daniel Keysers;Alexander Kolesnikov;Lucas Beyer
Xiaohua Zhai;Xiao Wang;Basil Mustafa;A. Steiner;Daniel Keysers;Alexander Kolesnikov;Lucas Beyer
中科院分区:
其他
文献类型:
--
作者:
Xiaohua Zhai;Xiao Wang;Basil Mustafa;A. Steiner;Daniel Keysers;Alexander Kolesnikov;Lucas Beyer

文献摘要

被引文献

相似文献

本文提出了对比调优,这是一种简单的方法,利用对比训练来对齐图像和文本模型,同时仍然利用它们的预训练。在我们的实证研究中,我们发现锁定的预训练图像模型与解锁的文本模型效果最好。我们称这种对比调优实例为“锁定图像调优”(LiT),它只是教文本模型从预训练的图像模型中读出良好的表示,以用于新任务。一个LiT模型获得了零射击转移到新的视觉任务的能力,如图像分类或检索。拟议的法律适用范围广泛;它可以可靠地与多种预训练方法(监督和非监督)以及使用三种不同的图像文本数据集的不同架构(ResNet, Vision Transformers和MLP-Mixer)一起工作。使用基于变压器的预训练ViT-g/14模型,LiT模型在ImageNet测试集上达到84.5%的零射击转移精度,在具有挑战性的非分布ObjectNet测试集上达到81.1%。
This paper presents contrastive-tuning, a simple method employing contrastive training to align image and text mod-els while still taking advantage of their pre-training. In our empirical study we find that locked pre-trained image mod-els with unlocked text models work best. We call this in-stance of contrastive-tuning “Locked-image Tuning” (LiT), which just teaches a text model to read out good repre-sentations from a pre-trained image model for new tasks. A LiT model gains the capability of zero-shot transfer to new vision tasks, such as image classification or retrieval. The proposed LiT is widely applicable; it works reliably with multiple pre-training methods (supervised and unsu-pervised) and across diverse architectures (ResNet, Vision Transformers and MLP-Mixer) using three different image-text datasets. With the transformer-based pre-trained ViT-g/14 model, the LiT model achieves 84.5% zero-shot trans-fer accuracy on the ImageNet test set, and 81.1% on the challenging out-of-distribution ObjectNet test set.