Positional relevance model for pseudo-relevance feedback

Positional relevance model for pseudo-relevance feedback
复制标题

DOI:
10.1145/1835449.1835546
复制
发表时间:
2010-07
期刊:
Proceedings of the 33rd international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
Yuanhua Lv;ChengXiang Zhai
Yuanhua Lv;ChengXiang Zhai
中科院分区:
其他
文献类型:
--
作者:
Yuanhua Lv;ChengXiang Zhai

文献摘要

被引文献

相似文献

伪相关反馈是改善检索结果的有效技术。传统的反馈算法使用整个反馈文档作为一个单元来提取查询扩展的单词,这不是最佳的,因为文档可能涵盖了几个不同的主题,因此包含许多无关的信息。在本文中,我们研究了如何根据反馈文档中的术语立场有效地从反馈文档中进行有效选择的单词。我们提出了一个位置相关模型(PRM),以统一的概率方式解决此问题。提出的PRM是相关模型的扩展,以利用术语位置和接近性,以便基于直觉将更接近查询单词的单词分配给单词更接近查询单词更可能与查询主题相关。我们开发了两种方法来基于不同的采样过程估算PRM。两个大检索数据集的实验结果表明,提出的PRM对于伪相关反馈是有效且健壮的,在基于文档的反馈和基于段落的反馈中都大大胜过相关模型。
Pseudo-relevance feedback is an effective technique for improving retrieval results. Traditional feedback algorithms use a whole feedback document as a unit to extract words for query expansion, which is not optimal as a document may cover several different topics and thus contain much irrelevant information. In this paper, we study how to effectively select from feedback documents those words that are focused on the query topic based on positions of terms in feedback documents. We propose a positional relevance model (PRM) to address this problem in a unified probabilistic way. The proposed PRM is an extension of the relevance model to exploit term positions and proximity so as to assign more weights to words closer to query words based on the intuition that words closer to query words are more likely to be related to the query topic. We develop two methods to estimate PRM based on different sampling processes. Experiment results on two large retrieval datasets show that the proposed PRM is effective and robust for pseudo-relevance feedback, significantly outperforming the relevance model in both document-based feedback and passage-based feedback.