Learning to Distribute Vocabulary Indexing for Scalable Visual Search

Learning to Distribute Vocabulary Indexing for Scalable Visual Search
复制标题

学习分配词汇索引以实现可扩展的视觉搜索

DOI:
10.1109/tmm.2012.2225035
复制
发表时间:
2013-01-01
影响因子:
7.3
通讯作者:
Gao, Wen
Gao, Wen
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ji, Rongrong;Duan, Ling-Yu;Gao, Wen

文献摘要

被引文献

相似文献

近些年来,基于倒排索引的近似重复视觉搜索范式的词袋搜索受到越来越多的关注。一个基本但尚未开发的挑战是如何在单个服务器内维护受其内存限制的大型索引结构,而内存限制极难扩展到数百万甚至数十亿个图像。在本文中,我们提出将近重复视觉搜索结构并行化,以索引多个服务器上的数百万幅图像,包括视觉词汇的分布和相应的索引结构。我们从机器学习的角度对词汇索引的分布进行了优化,提供了一种利用跨多个服务器的计算能力来减少搜索延迟的“记忆之光”搜索范例。特别是,我们的解决方案解决了两个基本问题:“分发什么”和“如何分发”。“分发什么”是通过“有损”词汇量增加来解决的,它在分发之前丢弃了频繁出现的和不加区别的单词。“如何分配”是通过学习最优分配函数来解决的,该最优分配函数最大化了将给定查询的词分配给多个服务器的一致性。我们在一个真实的位置搜索系统中对超过1000万幅地标图像进行了分布式词汇索引的验证。与目前最先进的单服务器搜索[5]、[6]、[16]和分布式搜索[23]相比,我们的方案在仅分发5%的单词的情况下,在相当的精度下获得了约200%的加速比。我们还报告说,即使在部分服务器崩溃的情况下,也具有出色的健壮性。
In recent years, there is an ever-increasing research focus on Bag-of-Words based near duplicate visual search paradigm with inverted indexing. One fundamental yet unexploited challenge is how to maintain the large indexing structures within a single server subject to its memory constraint, which is extremely hard to scale up to millions or even billions of images. In this paper, we propose to parallelize the near duplicate visual search architecture to index millions of images over multiple servers, including the distribution of both visual vocabulary and the corresponding indexing structure. We optimize the distribution of vocabulary indexing from a machine learning perspective, which provides a "memory light" search paradigm that leverages the computational power across multiple servers to reduce the search latency. Especially, our solution addresses two essential issues: "What to distribute" and "How to distribute". "What to distribute" is addressed by a "lossy" vocabulary Boosting, which discards both frequent and indiscriminating words prior to distribution. "How to distribute" is addressed by learning an optimal distribution function, which maximizes the uniformity of assigning the words of a given query to multiple servers. We validate the distributed vocabulary indexing scheme in a real world location search system over 10 million landmark images. Comparing to the state-of-the-art alternatives of single-server search [5], [6], [16] and distributed search [23], our scheme has yielded a significant gain of about 200% speedup at comparable precision by distributing only 5% words. We also report excellent robustness even when partial servers crash.