Efficient distributed selective search

Efficient distributed selective search
复制标题

DOI:
10.1007/s10791-016-9290-6
复制
发表时间:
2016-11
影响因子:
2.5
通讯作者:
Yubin Kim;Jamie Callan;J. Culpepper;Alistair Moffat
Yubin Kim;Jamie Callan;J. Culpepper;Alistair Moffat
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yubin Kim;Jamie Callan;J. Culpepper;Alistair Moffat

文献摘要

被引文献

相似文献

仿真和分析表明,选择性搜索可以降低大规模分布式信息检索的开销。 通过将集合划分为小的主题分片,然后使用资源排名算法选择分片的子集来搜索每个查询,可以评估更少的帖子。 在本文中,我们扩展到新的领域,使用细粒度的模拟研究选择性搜索,检查效率的差异时,基于术语和基于样本的资源选择算法使用;测量两个政策的影响分配索引碎片的机器;和探索的好处索引扩散和镜像部署的机器的数量是不同的。两个大的数据集和四个大的查询日志得到的结果证实,选择性搜索是显着更有效的比传统的分布式搜索架构,可以处理更高的查询率。此外,我们证明,选择性搜索可以调整,以避免瓶颈,从而最大限度地利用底层计算机硬件。
Simulation and analysis have shown thatselective searchcan reduce the cost of large-scale distributed information retrieval. By partitioning the collection into smalltopical shards, and then using a resource ranking algorithm to choose a subset of shards to search for each query, fewer postings are evaluated. In this paper we extend the study of selective search into new areas using a fine-grained simulation, examining the difference in efficiency when term-based and sample-based resource selection algorithms are used; measuring the effect of two policies for assigning index shards to machines; and exploring the benefits of index-spreading and mirroring as the number of deployed machines is varied. Results obtained for two large datasets and four large query logs confirm that selective search is significantly more efficient than conventional distributed search architectures and can handle higher query rates. Furthermore, we demonstrate that selective search can be tuned to avoid bottlenecks, and thus maximize usage of the underlying computer hardware.