Contrastive Learning with Complex Heterogeneity

Contrastive Learning with Complex Heterogeneity
复制标题

DOI:
10.1145/3534678.3539311
复制
发表时间:
2021-05
期刊:
Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining
影响因子:
--
通讯作者:
Lecheng Zheng;Jinjun Xiong;Yada Zhu;Jingrui He
Lecheng Zheng;Jinjun Xiong;Yada Zhu;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Lecheng Zheng;Jinjun Xiong;Yada Zhu;Jingrui He

文献摘要

被引文献

相似文献

随着跨多个高影响力应用程序的大数据的出现,我们经常面临复杂异构性的挑战。新收集的数据通常由多个模态组成,并具有多个标签的特征,从而表现出多种类型的异质性共存。尽管现有技术擅长于用足够的标签信息对复杂的异质性进行建模,但是在真实的应用中获得这样的标签信息可能是相当昂贵的。近年来,对比学习因其利用丰富的未标记数据的突出表现而受到研究者的关注。然而,现有的对比学习工作不能解决假阴性对的问题,即,如果某些“负”对具有相同的标记,则它们可能具有类似的表示。为了克服这些问题,在本文中,我们提出了一个统一的异构学习框架,它结合了加权无监督对比损失和加权监督对比损失来建模多种类型的异构性。我们首先提供了一个理论分析,表明香草对比学习损失很容易导致次优解决方案的假阴性对的存在,而建议的加权损失可以自动调整权重的基础上学习表示的相似性,以减轻这个问题。在真实数据集上的实验结果表明了该框架对多种异构性建模的有效性和效率。
With the advent of big data across multiple high-impact applications, we are often facing the challenge of complex heterogeneity. The newly collected data usually consist of multiple modalities and are characterized with multiple labels, thus exhibiting the co-existence of multiple types of heterogeneity. Although state-of-the-art techniques are good at modeling the complex heterogeneity with sufficient label information, such label information can be quite expensive to obtain in real applications. Recently, researchers pay great attention to contrastive learning due to its prominent performance by utilizing rich unlabeled data. However, existing work on contrastive learning is not able to address the problem of false-negative pairs, i.e., some 'negative' pairs may have similar representations if they have the same label. To overcome the issues, in this paper, we propose a unified heterogeneous learning framework, which combines both the weighted unsupervised contrastive loss and the weighted supervised contrastive loss to model multiple types of heterogeneity. We first provide a theoretical analysis showing that the vanilla contrastive learning loss easily leads to the sub-optimal solution in the presence of false-negative pairs, whereas the proposed weighted loss could automatically adjust the weight based on the similarity of the learned representations to mitigate this issue. Experimental results on real-world data sets demonstrate the effectiveness and the efficiency of the proposed framework modeling multiple types of heterogeneity.