DUET: Cross-modal Semantic Grounding for Contrastive Zero-shot Learning

DUET: Cross-modal Semantic Grounding for Contrastive Zero-shot Learning
复制标题

DOI:
10.48550/arxiv.2207.01328
复制
发表时间:
2022-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhuo Chen;Yufen Huang;Jiaoyan Chen;Yuxia Geng;Wen Zhang;Yin Fang;Jeff Z. Pan;Wenting Song;Huajun Chen
Zhuo Chen;Yufen Huang;Jiaoyan Chen;Yuxia Geng;Wen Zhang;Yin Fang;Jeff Z. Pan;Wenting Song;Huajun Chen
中科院分区:
其他
文献类型:
--
作者:
Zhuo Chen;Yufen Huang;Jiaoyan Chen;Yuxia Geng;Wen Zhang;Yin Fang;Jeff Z. Pan;Wenting Song;Huajun Chen

文献摘要

相似文献

Zero-shot learning(Zero-shot learning,简称ZRL)旨在预测在训练过程中从未出现过样本的不可见类。零炮图像分类中最有效和最广泛使用的语义信息之一是属性,它是类别级视觉特征的注释。然而,目前的方法往往无法区分这些微妙的视觉区别,不仅由于缺乏细粒度的注释,而且属性的不平衡和共生。在本文中,我们提出了一种基于transformer的端到端的语义学习方法DUET,它通过自监督多模态学习范式从预训练的语言模型(PLM)中集成潜在的语义知识。具体而言,我们(1)开发了一个跨模态语义基础网络,以研究模型的能力,从图像中分离语义属性;(2)应用属性级对比学习策略,以进一步提高模型的区分细粒度的视觉特征,对属性同现和不平衡;(3)提出了一个多任务的学习策略,考虑多模型的目标。我们发现,我们的DUET可以在三个标准的CNOL基准测试和一个配备知识图的CNOL基准测试上实现最先进的性能。其组成部分是有效的,其预测是可解释的。
Zero-shot learning (ZSL) aims to predict unseen classes whose samples have never appeared during training. One of the most effective and widely used semantic information for zero-shot image classification are attributes which are annotations for class-level visual characteristics. However, the current methods often fail to discriminate those subtle visual distinctions between images due to not only the shortage of fine-grained annotations, but also the attribute imbalance and co-occurrence. In this paper, we present a transformer-based end-to-end ZSL method named DUET, which integrates latent semantic knowledge from the pre-trained language models (PLMs) via a self-supervised multi-modal learning paradigm. Specifically, we (1) developed a cross-modal semantic grounding network to investigate the model's capability of disentangling semantic attributes from the images; (2) applied an attribute-level contrastive learning strategy to further enhance the model's discrimination on fine-grained visual characteristics against the attribute co-occurrence and imbalance; (3) proposed a multi-task learning policy for considering multi-model objectives. We find that our DUET can achieve state-of-the-art performance on three standard ZSL benchmarks and a knowledge graph equipped ZSL benchmark. Its components are effective and its predictions are interpretable.