Generic compact representation through visual-semantic ambiguity removal

Generic compact representation through visual-semantic ambiguity removal
复制标题

DOI:
10.1016/j.patrec.2018.04.024
复制
发表时间:
2019
期刊:
Pattern Recognit. Lett.
影响因子:
--
通讯作者:
Yang Long;Yu Guan;Ling Shao
Yang Long;Yu Guan;Ling Shao
中科院分区:
其他
文献类型:
--
作者:
Yang Long;Yu Guan;Ling Shao

文献摘要

相似文献

Zero-Shot Hashing(Zero-Shot Hashing)旨在学习紧凑的二进制代码,可以从看不见的类别中保留图像的语义内容。传统的方法将视觉特征投影到由可见和不可见类别共享的语义空间。然而,我们观察到,这样一个单向的范式遭受的视觉语义歧义问题。也就是说,语义概念(例如属性)不能明确地对应于视觉模式,反之亦然。这样的问题可能导致每个属性的视觉特征的巨大差异。在本文中,我们研究如何消除这种语义模糊的基础上观察到的视觉外观。具体而言,我们提出了(1)一种新的潜在属性空间,以弥补视觉外观和语义表达之间的差距;(2)一种双图正则化嵌入算法,称为视觉语义歧义消除(VSAR),它可以同时提取视觉和语义信息之间的共享组件,并根据两个空间的内在局部结构相互对齐数据分布;(3)提出了一种新的零次散列算法框架,该框架可以同时处理实例级和类别级任务。我们验证我们的方法在四个流行的基准。大量的实验表明,我们提出的方法显着执行国家的最先进的方法。
Zero-Shot Hashing (ZSH) aims to learn compact binary codes that can preserve semantic contents of the images from unseen categories. Conventional approaches project visual features to a semantic space that is shared by both seen and unseen categories. However, we observe that such a one-way paradigm suffers from thevisual-semantic ambiguityproblem. Namely, the semantic concepts (e.g. attributes) cannot explicitly correspond to visual patterns, and vice versa. Such a problem can lead to a huge variance in the visual features for each attribute. In this paper, we investigate how to remove such semantic ambiguity based on the observed visual appearances. In particular, we propose (1) a novel latent attribute space to mitigate the gap between visual appearances and semantic expressions; (2) a dual-graph regularised embedding algorithm calledVisual-SemanticAmbiguityRemoval(VSAR) that can simultaneously extract the shared components between visual and semantic information and mutually align the data distribution based on the intrinsic local structures of both spaces; (3) a new zero-shot hashing framework that can deal with both instance-level and category-level tasks. We validate our method on four popular benchmarks. Extensive experiments demonstrate that our proposed approach significantly performs the state-of-the-art methods.