Class-Weighted Convolutional Features for Visual Instance Search

Class-Weighted Convolutional Features for Visual Instance Search
复制标题

DOI:
10.5244/c.31.144
复制
发表时间:
2017-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Albert Jiménez;J. Álvarez;Xavier Giró-i-Nieto
Albert Jiménez;J. Álvarez;Xavier Giró-i-Nieto
中科院分区:
其他
文献类型:
--
作者:
Albert Jiménez;J. Álvarez;Xavier Giró-i-Nieto

文献摘要

被引文献

相似文献

在现实场景中,图像检索的目标是未标记图像的大型动态数据集。在这些情况下,每次向数据库中添加新图像时都对模型进行训练或微调既不高效,也不具有可扩展性。卷积神经网络训练图像分类在大数据集已被证明是有效的特征提取图像检索。最成功的方法是基于卷积层的激活编码,因为它们传递图像的空间信息。在本文中,我们超越了这种空间信息,提出了一种基于目标图像中预测的语义信息的卷积特征的局部感知编码。为此,我们使用类激活图(Class Activation Maps, CAMs)获得图像中最具判别性的区域。cam基于网络中包含的知识,因此,我们的方法具有不需要外部信息的额外优势。此外,我们使用CAMs在第一次快速搜索后的无监督重新排序阶段生成对象建议。我们在两个公共可用数据集Oxford5k和Paris6k上进行的实例检索实验表明,当使用在ImageNet上训练的现成模型时,我们的方法比当前最先进的方法更具竞争力。本文中使用的源代码和模型可以在这个http URL上公开获得。
Image retrieval in realistic scenarios targets large dynamic datasets of unlabeled images. In these cases, training or fine-tuning a model every time new images are added to the database is neither efficient nor scalable. Convolutional neural networks trained for image classification over large datasets have been proven effective feature extractors for image retrieval. The most successful approaches are based on encoding the activations of convolutional layers, as they convey the image spatial information. In this paper, we go beyond this spatial information and propose a local-aware encoding of convolutional features based on semantic information predicted in the target image. To this end, we obtain the most discriminative regions of an image using Class Activation Maps (CAMs). CAMs are based on the knowledge contained in the network and therefore, our approach, has the additional advantage of not requiring external information. In addition, we use CAMs to generate object proposals during an unsupervised re-ranking stage after a first fast search. Our experiments on two public available datasets for instance retrieval, Oxford5k and Paris6k, demonstrate the competitiveness of our approach outperforming the current state-of-the-art when using off-the-shelf models trained on ImageNet. The source code and model used in this paper are publicly available at this http URL.