Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation Learning

Synthetic Data Can Also Teach: Synthesizing Effective Data for Unsupervised Visual Representation Learning
复制标题

DOI:
10.1609/aaai.v37i3.25388
复制
发表时间:
2022-02
期刊:
--
影响因子:
--
通讯作者:
Yawen Wu;Zhepeng Wang;Dewen Zeng;Yiyu Shi;Jingtong Hu
Yawen Wu;Zhepeng Wang;Dewen Zeng;Yiyu Shi;Jingtong Hu
中科院分区:
其他
文献类型:
--
作者:
Yawen Wu;Zhepeng Wang;Dewen Zeng;Yiyu Shi;Jingtong Hu

文献摘要

被引文献

相似文献

对比学习(CL)是一种自监督学习方法,能够有效地从未标记数据中学习视觉表征。给定对比学习的训练数据,可以训练生成模型来生成合成数据以补充真实数据。同时使用合成数据和真实数据进行对比学习训练有可能提高所学表征的质量。然而,合成数据通常质量低于真实数据,并且与使用真实数据相比,使用合成数据可能无法改进对比学习。为了解决这个问题,我们提出了一个数据生成框架,其中包含两种通过联合样本生成和对比学习来改进对比学习训练的方法。第一种方法是为主要模型生成困难样本。生成器与主要模型联合学习,以便根据主要模型的训练状态动态定制困难样本。此外,还提出了一对数据生成器来生成相似但不同的样本作为正对。在联合学习中,通过降低正对的相似性来逐步增加其难度。在多个数据集上的实验结果表明,应用于对比学习的所提出的数据生成方法具有更高的准确性和数据效率。例如,在ImageNet - 100、CIFAR - 100和CIFAR - 10上,线性分类的准确率分别提高了约4.0%、3.5%和2.6%。此外,线性分类的数据效率提高了多达2倍,迁移学习的数据效率提高了多达5倍。
Contrastive learning (CL), a self-supervised learning approach, can effectively learn visual representations from unlabeled data. Given the CL training data, generative models can be trained to generate synthetic data to supplement the real data. Using both synthetic and real data for CL training has the potential to improve the quality of learned representations. However, synthetic data usually has lower quality than real data, and using synthetic data may not improve CL compared with using real data. To tackle this problem, we propose a data generation framework with two methods to improve CL training by joint sample generation and contrastive learning. The first approach generates hard samples for the main model. The generator is jointly learned with the main model to dynamically customize hard samples based on the training state of the main model. Besides, a pair of data generators are proposed to generate similar but distinct samples as positive pairs. In joint learning, the hardness of a positive pair is progressively increased by decreasing their similarity. Experimental results on multiple datasets show superior accuracy and data efficiency of the proposed data generation methods applied to CL. For example, about 4.0%, 3.5%, and 2.6% accuracy improvements for linear classification are observed on ImageNet-100, CIFAR-100, and CIFAR-10, respectively. Besides, up to 2× data efficiency for linear classification and up to 5× data efficiency for transfer learning are achieved.