Instance Adaptive Prototypical Contrastive Embedding for Generalized Zero Shot Learning

Instance Adaptive Prototypical Contrastive Embedding for Generalized Zero Shot Learning
复制标题

DOI:
10.48550/arxiv.2309.06987
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Riti Paul;Sahil Vora;Baoxin Li
Riti Paul;Sahil Vora;Baoxin Li
中科院分区:
其他
文献类型:
--
作者:
Riti Paul;Sahil Vora;Baoxin Li

文献摘要

相似文献

假设在训练过程中无法访问看不见的标签,则旨在从可见和看不见的标签中对样本进行分类。 GZSL的最新进展通过将基于对比的(基于实例的)嵌入生成网络中并利用数据点之间的语义关系加快。但是,现有的嵌入体系结构遭受了两个局限性:(1)合成特征嵌入的有限可区分性,而无需考虑细粒的簇结构; (2)由于现有的对比度嵌入网络的限制缩放机制而引起的不灵活的优化,导致嵌入空间中的代表性重叠。为了提高(1)中提到的嵌入空间中表示的质量,我们提出了一个基于保证金的原型对比度学习嵌入网络,从在为嵌入网络和生成器提供大量的集群监督的同时,粒度表示)交互。为了解决(2),我们提出了一个实例自适应对比损失,从而导致跨阶层间边缘增加的看不见标签的普遍表示。通过全面的实验评估,我们表明我们的方法可以胜过三个基准数据集的当前最新技术。我们的方法还始终在GZSL环境中实现最好的看不见的性能。
Generalized zero-shot learning(GZSL) aims to classify samples from seen and unseen labels, assuming unseen labels are not accessible during training. Recent advancements in GZSL have been expedited by incorporating contrastive-learning-based (instance-based) embedding in generative networks and leveraging the semantic relationship between data points. However, existing embedding architectures suffer from two limitations: (1) limited discriminability of synthetic features' embedding without considering fine-grained cluster structures; (2) inflexible optimization due to restricted scaling mechanisms on existing contrastive embedding networks, leading to overlapped representations in the embedding space. To enhance the quality of representations in the embedding space, as mentioned in (1), we propose a margin-based prototypical contrastive learning embedding network that reaps the benefits of prototype-data (cluster quality enhancement) and implicit data-data (fine-grained representations) interaction while providing substantial cluster supervision to the embedding network and the generator. To tackle (2), we propose an instance adaptive contrastive loss that leads to generalized representations for unseen labels with increased inter-class margin. Through comprehensive experimental evaluation, we show that our method can outperform the current state-of-the-art on three benchmark datasets. Our approach also consistently achieves the best unseen performance in the GZSL setting.