Toward a higher-level visual representation for object-based image retrieval

Toward a higher-level visual representation for object-based image retrieval
复制标题

DOI:
10.1007/s00371-008-0294-0
复制
发表时间:
2008-11
期刊:
The Visual Computer
影响因子:
--
通讯作者:
Yantao Zheng;Shi-Yong Neo;Tat-Seng Chua;Q. Tian
Yantao Zheng;Shi-Yong Neo;Tat-Seng Chua;Q. Tian
中科院分区:
其他
文献类型:
--
作者:
Yantao Zheng;Shi-Yong Neo;Tat-Seng Chua;Q. Tian

文献摘要

被引文献

相似文献

我们提出了一种更高级别的视觉表示,视觉同义词集,用于超越视觉外观的基于对象的图像检索。所提出的视觉表示在两个方面改进了传统的基于部分的词袋图像表示。首先,该方法通过从频繁共现的视觉词集中构建中间描述符(视觉短语)来增强视觉词的辨别能力。其次,为了弥合视觉外观差异或实现更好的类内不变性能力,该方法根据视觉单词和短语的类概率分布将其聚类为视觉同义词集。基本原理是视觉单词或短语的分布往往在其所属对象类周围达到峰值。在Caltech-256数据集上的测试表明,视觉同义词集可以部分弥合同一类图像的视觉差异,并对具有不同视觉外观的相关图像提供令人满意的检索。
We propose a higher-level visual representation, visual synset, for object-based image retrieval beyond visual appearances. The proposed visual representation improves the traditional part-based bag-of-words image representation, in two aspects. First, the approach strengthens the discrimination power of visual words by constructing an intermediate descriptor, visual phrase, from frequently co-occurring visual word-set. Second, to bridge the visual appearance difference or to achieve better intra-class invariance power, the approach clusters visual words and phrases into visual synset, based on their class probability distribution. The rationale is that the distribution of visual word or phrase tends to peak around its belonging object classes. The testing on Caltech-256 data set shows that the visual synset can partially bridge visual differences of images of the same class and deliver satisfactory retrieval of relevant images with different visual appearances.