Towards Affordable Semantic Searching: Zero-Shot Retrieval via Dominant Attributes

Towards Affordable Semantic Searching: Zero-Shot Retrieval via Dominant Attributes
复制标题

DOI:
10.1609/aaai.v32i1.12280
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
Yang Long;Li Liu;Yuming Shen;Ling Shao
Yang Long;Li Liu;Yuming Shen;Ling Shao
中科院分区:
其他
文献类型:
--
作者:
Yang Long;Li Liu;Yuming Shen;Ling Shao

文献摘要

被引文献

相似文献

实例级检索已经成为从大规模数据库中索引和检索图像的基本范例。传统的实例搜索需要查询图像的至少一个示例来检索包含相同对象实例的图像。现有的语义检索只能搜索语义相关的图像,如共享相同类别或一组标签的图像,而不是确切的实例。同时,不切实际的假设是,所有类别或标签都是事先知道的。这些语义概念的训练模型高度依赖实例级属性或人工字幕,而这些属性或字幕的获取成本很高。考虑到上述挑战,本文研究了仅使用几个主要属性进行实例级图像搜索的零镜头检索问题。本文的主要贡献是:1)利用词的自动嵌入来推断类级属性,避免了昂贵的人工标注;2)通过我们提出的潜在实例属性发现(LIAD)算法,可以将推断出的类属性扩展为区分实例属性;3)我们的方法不仅限于完整的属性签名,还可以处理显性属性的查询。在CUB和SUN两个基准测试上的大量实验表明,我们的方法可以获得令人满意的性能。此外,我们的方法也可以使传统的ZSL任务受益。
Instance-level retrieval has become an essential paradigm to index and retrieves images from large-scale databases. Conventional instance search requires at least an example of the query image to retrieve images that contain the same object instance. Existing semantic retrieval can only search semantically-related images, such as those sharing the same category or a set of tags, not the exact instances. Meanwhile, the unrealistic assumption is that all categories or tags are known beforehand. Training models for these semantic concepts highly rely on instance-level attributes or human captions which are expensive to acquire. Given the above challenges, this paper studies the Zero-shot Retrieval problem that aims for instance-level image search using only a few dominant attributes. The contributions are: 1) we utilise automatic word embedding to infer class-level attributes to circumvent expensive human labelling; 2) the inferred class-attributes can be extended into discriminative instance attributes through our proposed Latent Instance Attributes Discovery (LIAD) algorithm; 3) our method is not restricted to complete attribute signatures, query of dominant attributes can also be dealt with. On two benchmarks, CUB and SUN, extensive experiments demonstrate that our method can achieve promising performance for the problem. Moreover, our approach can also benefit conventional ZSL tasks.