Can contrastive learning avoid shortcut solutions?

Can contrastive learning avoid shortcut solutions?
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
Advances in neural information processing systems
影响因子:
--
通讯作者:
Joshua Robinson;Li Sun;Ke Yu;K. Batmanghelich;S. Jegelka;S. Sra
Joshua Robinson;Li Sun;Ke Yu;K. Batmanghelich;S. Jegelka;S. Sra
中科院分区:
其他
文献类型:
--
作者:
Joshua Robinson;Li Sun;Ke Yu;K. Batmanghelich;S. Jegelka;S. Sra

文献摘要

相似文献

通过对比学习学习到的表示的泛化很大程度上取决于提取的数据特征。然而,我们观察到对比损失并不总是充分指导提取哪些特征,这种行为可能会通过“捷径”(即无意中抑制重要的预测特征)对下游任务的性能产生负面影响。我们发现特征提取受到所谓的实例辨别任务(即区分相似点对和不相似点对的任务)难度的影响。尽管较难的对改进了某些特征的表示,但这种改进是以抑制先前良好表示的特征为代价的。为此,我们提出了隐式特征修改(IFM),这是一种改变正样本和负样本的方法,以指导对比模型捕获更广泛的预测特征。根据经验,我们观察到 IFM 减少了特征抑制,从而提高了视觉和医学成像任务的性能。该代码位于:https://github.com/joshr17/IFM。
The generalization of representations learned via contrastive learning depends crucially on what features of the data are extracted. However, we observe that the contrastive loss does not always sufficiently guide which features are extracted, a behavior that can negatively impact the performance on downstream tasks via "shortcuts", i.e., by inadvertently suppressing important predictive features. We find that feature extraction is influenced by the difficulty of the so-called instance discrimination task (i.e., the task of discriminating pairs of similar points from pairs of dissimilar ones). Although harder pairs improve the representation of some features, the improvement comes at the cost of suppressing previously well represented features. In response, we propose implicit feature modification (IFM), a method for altering positive and negative samples in order to guide contrastive models towards capturing a wider variety of predictive features. Empirically, we observe that IFM reduces feature suppression, and as a result improves performance on vision and medical imaging tasks. The code is available at: https://github.com/joshr17/IFM.