Self-supervised Learning from a Multi-view Perspective

Self-supervised Learning from a Multi-view Perspective
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Yao-Hung Hubert Tsai;Yue Wu;R. Salakhutdinov;Louis-Philippe Morency
Yao-Hung Hubert Tsai;Yue Wu;R. Salakhutdinov;Louis-Philippe Morency
中科院分区:
其他
文献类型:
--
作者:
Yao-Hung Hubert Tsai;Yue Wu;R. Salakhutdinov;Louis-Philippe Morency

文献摘要

被引文献

相似文献

作为无监督表示学习的一个子集,自监督表示学习采用自定义信号作为监督,并将学习到的表示用于下游任务,例如对象检测和图像字幕。许多提出的自监督学习方法自然遵循多视图视角,其中输入(例如原始图像)和自监督信号(例如增强图像)可以被视为数据的两个冗余视图。从这种多视角的角度出发,本文提供了一个信息理论框架,以更好地理解鼓励成功自我监督学习的属性。具体来说,我们证明自监督学习表示可以提取与任务相关的信息并丢弃与任务无关的信息。我们的理论框架为更大的自我监督学习目标设计空间铺平了道路。特别是,我们提出了一个复合目标,弥合了先前的对比学习目标和预测学习目标之间的差距,并引入了一个额外的目标术语来丢弃与任务无关的信息。为了验证我们的分析,我们进行了对照实验来评估复合目标的影响。我们还探索了我们的框架在多视图视角之外的经验概括,其中跨视图冗余可能无法被清楚地观察到。
As a subset of unsupervised representation learning, self-supervised representation learning adopts self-defined signals as supervision and uses the learned representation for downstream tasks, such as object detection and image captioning. Many proposed approaches for self-supervised learning follow naturally a multi-view perspective, where the input (e.g., original images) and the self-supervised signals (e.g., augmented images) can be seen as two redundant views of the data. Building from this multi-view perspective, this paper provides an information-theoretical framework to better understand the properties that encourage successful self-supervised learning. Specifically, we demonstrate that self-supervised learned representations can extract task-relevant information and discard task-irrelevant information. Our theoretical framework paves the way to a larger space of self-supervised learning objective design. In particular, we propose a composite objective that bridges the gap between prior contrastive and predictive learning objectives, and introduce an additional objective term to discard task-irrelevant information. To verify our analysis, we conduct controlled experiments to evaluate the impact of the composite objectives. We also explore our framework's empirical generalization beyond the multi-view perspective, where the cross-view redundancy may not be clearly observed.