Efficiency Implications of Term Weighting for Passage Retrieval

Efficiency Implications of Term Weighting for Passage Retrieval
复制标题

DOI:
10.1145/3397271.3401263
复制
发表时间:
2020-07
期刊:
Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
J. Mackenzie;Zhuyun Dai;L. Gallagher;Jamie Callan
J. Mackenzie;Zhuyun Dai;L. Gallagher;Jamie Callan
中科院分区:
其他
文献类型:
--
作者:
J. Mackenzie;Zhuyun Dai;L. Gallagher;Jamie Callan

文献摘要

被引文献

相似文献

语言模型预训练已经引起了人们对涉及自然语言理解的任务的极大关注,并已成功应用于许多下游任务,取得了令人印象深刻的结果。在信息检索中,许多解决方案成本太高,无法独立存在,需要多级排名架构。最近的工作已经开始考虑如何“backport”这些计算昂贵的模型的显着方面的检索管道的前几个阶段。其中一个例子是DeepCT,它使用BERT在段落级别的给定上下文中重新加权术语重要性。这个过程是离线计算的,会产生一个具有重新加权词频值的增强倒排索引。在这项工作中,我们对DeepCT索引的查询处理效率进行了调查。使用一些候选生成算法,我们揭示了术语重新加权如何影响查询处理延迟,并探索了如何将DeepCT用作静态索引修剪技术来加速查询处理而不损害搜索效率。
Language model pre-training has spurred a great deal of attention for tasks involving natural language understanding, and has been successfully applied to many downstream tasks with impressive results. Within information retrieval, many of these solutions are too costly to stand on their own, requiring multi-stage ranking architectures. Recent work has begun to consider how to "backport" salient aspects of these computationally expensive models to previous stages of the retrieval pipeline. One such instance is DeepCT, which uses BERT to re-weight term importance in a given context at the passage level. This process, which is computed offline, results in an augmented inverted index with re-weighted term frequency values. In this work, we conduct an investigation of query processing efficiency over DeepCT indexes. Using a number of candidate generation algorithms, we reveal how term re-weighting can impact query processing latency, and explore how DeepCT can be used as a static index pruning technique to accelerate query processing without harming search effectiveness.