More Synergy, Less Redundancy: Exploiting Joint Mutual Information for Self-Supervised Learning

More Synergy, Less Redundancy: Exploiting Joint Mutual Information for Self-Supervised Learning
复制标题

DOI:
10.1109/icip49359.2023.10222547
复制
发表时间:
2023-07
期刊:
2023 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
S. Mohamadi;Gianfranco Doretto;Don Adjeroh
S. Mohamadi;Gianfranco Doretto;Don Adjeroh
中科院分区:
其他
文献类型:
--
作者:
S. Mohamadi;Gianfranco Doretto;Don Adjeroh

文献摘要

相似文献

自监督学习(SSL)现在是监督学习的重要竞争对手,尽管它不需要数据注释。几个基线试图使SSL模型利用有关数据分布的信息,并减少对增强效应的依赖。然而,有没有明确的共识,最大化或最小化增强视图的表示之间的互信息实际上有助于提高或降低SSL模型的性能。本文是一个基础性的工作,我们调查的作用,在SSL的互信息,并重新制定的SSL问题的背景下,一个新的角度对互信息。为此,我们认为联合互信息的角度来看,部分信息分解(PID)作为一个关键步骤,在可靠的多元信息测量。PID使我们能够将联合互信息分解为三个重要组成部分,即唯一信息,冗余信息和协同信息。我们的框架旨在最大限度地减少视图和所需的目标表示之间的冗余信息,同时最大限度地提高协同信息。我们的实验导致两个冗余减少基线的重新校准,并提出了一个新的SSL训练协议。在多个数据集和两个下游任务上的实验结果表明了该框架的有效性。
Self-supervised learning (SSL) is now a serious competitor for supervised learning, even though it does not require data annotation. Several baselines have attempted to make SSL models exploit information about data distribution, and less dependent on the augmentation effect. However, there is no clear consensus on whether maximizing or minimizing the mutual information between representations of augmentation views practically contribute to improvement or degradation in performance of SSL models. This paper is a fundamental work where, we investigate the role of mutual information in SSL, and reformulate the problem of SSL in the context of a new perspective on mutual information. To this end, we consider joint mutual information from the perspective of partial information decomposition (PID) as a key step in reliable multivariate information measurement. PID enables us to decompose joint mutual information into three important components, namely, unique information, redundant information and synergistic information. Our framework aims for minimizing the redundant information between views and the desired target representation while maximizing the synergistic information at the same time. Our experiments lead to a re-calibration of two redundancy reduction baselines, and a proposal for a new SSL training protocol. Experimental results on multiple datasets and two downstream tasks show the effectiveness of this framework.