The Trade-off between Universality and Label Efficiency of Representations from Contrastive Learning

The Trade-off between Universality and Label Efficiency of Representations from Contrastive Learning
复制标题

DOI:
10.48550/arxiv.2303.00106
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhenmei Shi;Jiefeng Chen;Kunyang Li;Jayaram Raghuram;Xi Wu;Yingyu Liang;S. Jha
Zhenmei Shi;Jiefeng Chen;Kunyang Li;Jayaram Raghuram;Xi Wu;Yingyu Liang;S. Jha
中科院分区:
其他
文献类型:
--
作者:
Zhenmei Shi;Jiefeng Chen;Kunyang Li;Jayaram Raghuram;Xi Wu;Yingyu Liang;S. Jha

文献摘要

相似文献

预训练表征(又名基础模型)最近已成为一种流行的学习范式,其中首先使用大规模无标签数据预训练一种表征,然后使用来自下游任务的少量有标签数据在该表征之上学习简单的预测器。对于该表征有两个关键的期望特性:标签效率(使用少量有标签数据在表征之上学习准确分类器的能力)和通用性(在广泛的下游任务中的有用性)。在本文中,我们聚焦于这种范式最流行的实例之一:带有线性探测的对比学习,即对通过对比学习预训练的表征学习一个线性预测器。我们表明在这两个期望特性之间存在一种权衡,以至于可能无法同时实现两者。具体而言,我们使用一种理论数据模型进行分析并表明,虽然更多样化的预训练数据会为不同任务产生更多样化的特征(提高通用性),但它对特定任务的特征关注较少,导致下游监督任务的样本复杂度更高,从而预测性能更差。在这一分析的指导下,我们提出一种对比正则化方法来改善这种权衡。我们使用真实世界的数据集和基础模型通过系统实验从经验上验证了我们的分析和方法。
Pre-training representations (a.k.a. foundation models) has recently become a prevalent learning paradigm, where one first pre-trains a representation using large-scale unlabeled data, and then learns simple predictors on top of the representation using small labeled data from the downstream tasks. There are two key desiderata for the representation: label efficiency (the ability to learn an accurate classifier on top of the representation with a small amount of labeled data) and universality (usefulness across a wide range of downstream tasks). In this paper, we focus on one of the most popular instantiations of this paradigm: contrastive learning with linear probing, i.e., learning a linear predictor on the representation pre-trained by contrastive learning. We show that there exists a trade-off between the two desiderata so that one may not be able to achieve both simultaneously. Specifically, we provide analysis using a theoretical data model and show that, while more diverse pre-training data result in more diverse features for different tasks (improving universality), it puts less emphasis on task-specific features, giving rise to larger sample complexity for down-stream supervised tasks, and thus worse prediction performance. Guided by this analysis, we propose a contrastive regularization method to improve the trade-off. We validate our analysis and method empirically with systematic experiments using real-world datasets and foundation models.