Debiased Contrastive Learning

Debiased Contrastive Learning
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Ching-Yao Chuang;Joshua Robinson;Yen-Chen Lin;A. Torralba;S. Jegelka
Ching-Yao Chuang;Joshua Robinson;Yen-Chen Lin;A. Torralba;S. Jegelka
中科院分区:
其他
文献类型:
--
作者:
Ching-Yao Chuang;Joshua Robinson;Yen-Chen Lin;A. Torralba;S. Jegelka

文献摘要

被引文献

相似文献

自监督表示学习的一项突出技术是对比语义相似和不相似的样本对。如果无法访问标签,不同的(负)点通常被视为随机采样的数据点,隐含地接受这些点实际上可能具有相同的标签。也许并不奇怪,我们观察到,在标签可用的合成环境中,从真正不同的标签中采样负面示例可以提高性能。受这一观察的启发,我们开发了一个去偏对比目标,即使在不了解真实标签的情况下,也可以纠正相同标签数据点的采样。根据经验,所提出的目标在视觉、语言和强化学习基准方面始终优于表征学习的最新技术。理论上,我们为下游分类任务建立泛化界限。
A prominent technique for self-supervised representation learning has been to contrast semantically similar and dissimilar pairs of samples. Without access to labels, dissimilar (negative) points are typically taken to be randomly sampled datapoints, implicitly accepting that these points may, in reality, actually have the same label. Perhaps unsurprisingly, we observe that sampling negative examples from truly different labels improves performance, in a synthetic setting where labels are available. Motivated by this observation, we develop a debiased contrastive objective that corrects for the sampling of same-label datapoints, even without knowledge of the true labels. Empirically, the proposed objective consistently outperforms the state-of-the-art for representation learning in vision, language, and reinforcement learning benchmarks. Theoretically, we establish generalization bounds for the downstream classification task.