Incorporating Contextual Information into Bag-of-Visual-Words Framework for Effective Object Categorization

Incorporating Contextual Information into Bag-of-Visual-Words Framework for Effective Object Categorization
复制标题

DOI:
10.1587/transinf.e95.d.3060
复制
发表时间:
2012-12
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Shuang Bai;Tetsuya Matsumoto;Y. Takeuchi;H. Kudo;N. Ohnishi
Shuang Bai;Tetsuya Matsumoto;Y. Takeuchi;H. Kudo;N. Ohnishi
中科院分区:
其他
文献类型:
--
作者:
Shuang Bai;Tetsuya Matsumoto;Y. Takeuchi;H. Kudo;N. Ohnishi

文献摘要

相似文献

视觉词袋是一种很有前途的物体分类方法。然而,在该框架中,由于矢量量化导致的信息丢失,视觉单词的补丁编码存在歧义。在本文中,我们提出将补丁级别的上下文信息加入到视觉词包中,以减少上述歧义。为了实现这一目标,我们构造了一种分层码本,其中上层的视觉单词包含下层视觉单词的上下文信息。在该方法中,我们从每个样本点提取不同尺度的斑块,所有这些斑块都用SIFT描述符来描述。然后,我们构建了分层码本,其中由粗尺度块产生的视觉单词被放在上层,而由细尺度块产生的视觉单词被放在下层。同时,利用提取出的这些块之间的对应关系,将不同层次的视觉单词关联起来。然后,我们设计了一种方法,将从同一采样点提取的补丁对分配到构建的码本上。此外,为了有效地利用图像信息,我们实现了基于不同采样策略提取的两组特征并使用概率方法进行融合的方法。最后,我们在数据集Caltech 101和数据集Caltech 256上对所提出的方法进行了评估。实验结果证明了该方法的有效性。关键词:物体分类、视觉词袋、语境信息、分层码本
Bag of visual words is a promising approach to object categorization. However, in this framework, ambiguity exists in patch encoding by visual words, due to information loss caused by vector quantization. In this paper, we propose to incorporate patch-level contextual information into bag of visual words for reducing the ambiguity mentioned above. To achieve this goal, we construct a hierarchical codebook in which visual words in the upper hierarchy contain contextual information of visual words in the lower hierarchy. In the proposed method, from each sample point we extract patches of different scales, all of which are described by the SIFT descriptor. Then, we build the hierarchical codebook in which visual words created from coarse scale patches are put in the upper hierarchy, while visual words created from fine scale patches are put in the lower hierarchy. At the same time, by employing the corresponding relationship among these extracted patches, visual words in different hierarchies are associated with each other. After that, we design a method to assign patch pairs, whose patches are extracted from the same sample point, to the constructed codebook. Furthermore, to utilize image information effectively, we implement the proposed method based on two sets of features which are extracted through different sampling strategies and fuse them using a probabilistic approach. Finally, we evaluate the proposed method on dataset Caltech 101 and dataset Caltech 256. Experimental results demonstrate the effectiveness of the proposed method. key words: object categorization, bag of visual words, contextual information, hierarchical codebook