Aggregating Local Image Descriptors into Compact Codes

Aggregating Local Image Descriptors into Compact Codes
复制标题

DOI:
10.1109/tpami.2011.235
复制
发表时间:
2012-09-01
影响因子:
23.6
通讯作者:
Schmid, Cordelia
Schmid, Cordelia
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jegou, Herve;Perronnin, Florent;Schmid, Cordelia

文献摘要

被引文献

相似文献

本文探讨了大规模图像搜索的问题。必须考虑三个约束条件:搜索准确性、效率和内存使用。我们首先提出并评估了将局部图像描述符聚合为向量的不同方法,并表明对于任何给定的向量维度,费舍尔核(Fisher kernel)都比参考的视觉词袋(bag-of-visual words)方法取得更好的性能。然后我们联合优化降维和索引,以获得精确的向量比较以及紧凑的表示。评估表明,图像表示可以减少到几十字节,同时保持高精度。在一个处理器核心上搜索一亿张图像的数据集大约需要250毫秒。
This paper addresses the problem of large-scale image search. Three constraints have to be taken into account: search accuracy, efficiency, and memory usage. We first present and evaluate different ways of aggregating local image descriptors into a vector and show that the Fisher kernel achieves better performance than the reference bag-of-visual words approach for any given vector dimension. We then jointly optimize dimensionality reduction and indexing in order to obtain a precise vector comparison as well as a compact representation. The evaluation shows that the image representation can be reduced to a few dozen bytes while preserving high accuracy. Searching a 100 million image data set takes about 250 ms on one processor core.