Investigating the Bilateral Connections in Generative Zero-Shot Learning

Investigating the Bilateral Connections in Generative Zero-Shot Learning
复制标题

DOI:
10.1109/tcyb.2021.3050803
复制
发表时间:
2021-02
影响因子:
11.8
通讯作者:
Jingjing Li;Mengmeng Jing;Ke Lu;Lei Zhu;Heng Tao Shen
Jingjing Li;Mengmeng Jing;Ke Lu;Lei Zhu;Heng Tao Shen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jingjing Li;Mengmeng Jing;Ke Lu;Lei Zhu;Heng Tao Shen

文献摘要

被引文献

相似文献

零射击学习(Zero-Shot Learning,简称ZRL)是计算机视觉领域一个非常有趣的话题,因为它处理新的实例和看不见的类别。在一个典型的语境中,存在着一个主要的视觉空间和一个辅助的语义空间。大多数现有的语义学习方法通过学习视觉到语义的映射或语义到视觉的映射来处理这个问题。换句话说,他们研究的是从一端到另一端的单向联系。然而,视觉空间与语义空间之间的联系实际上是双向的,即视觉空间描绘语义空间,语义空间描述视觉空间。因此,在这篇文章中,我们研究了BNL中的双边连接,并通过利用条件生成对抗网络(GAN)提出了一种新的模型,称为Boomerang-GAN。具体来说,我们通过条件GAN从其类别语义嵌入中生成看不见的视觉样本。与现有的生成式语义特征提取方法只考虑从类描述中生成视觉特征不同,我们的方法还考虑通过引入多模态循环一致性损失,将生成的视觉特征翻译回相应的语义嵌入.在五个广泛使用的数据集上进行的大量实验验证了我们的方法在识别和分割任务中都能够优于以前的最先进方法。
Zero-shot learning (ZSL) is a pretty intriguing topic in the computer vision community since it handles novel instances and unseen categories. In a typical ZSL setting, there is a main visual space and an auxiliary semantic space. Most existing ZSL methods handle the problem by learning either a visual-to-semantic mapping or a semantic-to-visual mapping. In other words, they investigate a unilateral connection from one end to the other. However, the connection between the visual space and the semantic space are bilateral in reality, that is, the visual space depicts the semantic space; the semantic space, on the other hand, describes the visual space. In this article, therefore, we investigate the bilateral connections in ZSL and present a novel model, called Boomerang-GAN, by taking advantage of conditional generative adversarial networks (GANs). Specifically, we generate unseen visual samples from their category semantic embeddings by a conditional GAN. Different from the existing generative ZSL methods that only consider generating visual features from class descriptions, our method also considers that the generated visual features can be translated back to their corresponding semantic embeddings by introducing a multimodal cycle-consistent loss. Extensive experiments of both ZSL and generalized ZSL on five widely used datasets verify that our method is able to outperform previous state-of-the-art approaches in both recognition and segmentation tasks.