A metric cache for similarity search

A metric cache for similarity search
复制标题

用于相似性搜索的度量缓存

DOI:
--
复制
发表时间:
2008
期刊:
Large-Scale Distributed Systems for Information Retrieval
影响因子:
--
通讯作者:
F. Rabitti
F. Rabitti
中科院分区:
--
文献类型:
--
作者:
F. Falchi;C. Lucchese;S. Orlando;R. Perego;F. Rabitti

文献摘要

被引文献

相似文献

度量空间中的相似性搜索是一个通用的范例,可以用于许多应用领域。它也可以有效地利用基于内容的图像检索系统,这是他们的目标转向网络规模的尺寸。在这种情况下,一个重要的问题是设计可伸缩的解决方案,它结合了联合收割机并行和分布式架构与高速缓存在几个级别。 为此,我们调查的相似性缓存,在度量空间的设计。它能够回答精确和近似的结果:即使缓存中不存在精确匹配,我们的缓存也可能返回具有质量保证的近似结果集。通过对100万张高质量的数码照片进行测试,我们表明,所提出的缓存技术可以对性能产生显着的影响,就像对传统Web搜索引擎的文本查询缓存已被证明是有效的。
Similarity search in metric spaces is a general paradigm that can be used in several application fields. It can also be effectively exploited in content-based image retrieval systems, which are shifting their target towards the Web-scale dimension. In this context, an important issue becomes the design of scalable solutions, which combine parallel and distributed architectures with caching at several levels. To this end, we investigate the design of a similarity cache that works in metric spaces. It is able to answer with exact and approximate results: even when an exact match is not present in cache, our cache may return an approximate result set with quality guarantees. By conducting tests on a collection of one million high-quality digital photos, we show that the proposed caching techniques can have a significant impact on performance, like caching on text queries has been proved effective for traditional Web search engines.