Automatic query wefinement using lexical affinities with maximal information gain

Automatic query wefinement using lexical affinities with maximal information gain
复制标题

使用具有最大信息增益的词汇亲和力自动查询优化

DOI:
--
复制
发表时间:
2002
期刊:
Annual International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
A. Soffer
A. Soffer
中科院分区:
--
文献类型:
--
作者:
David Carmel;E. Farchi;Yael Petruschka;A. Soffer

文献摘要

被引文献

相似文献

这项工作描述了一个自动查询细化技术,其重点是提高精度的排名靠前的文件。用于细化的术语是词汇亲和度(LA),即精确包含原始查询术语之一的密切相关的词对。将这些术语添加到查询中相当于对搜索结果重新排序,因此,在保留召回率的同时提高了精确度。我们描述了一种新的方法,选择最“翔实”的LAs细化,即那些LAs,最好的分离相关文件从不相关的文件中的结果。使用基于搜索引擎的评分函数的无监督估计来确定候选LA的信息增益。因此,该方法是全自动的,其质量取决于评分函数的质量。我们用TREC数据进行的实验清楚地表明,排名靠前的文档的精度有了显着提高。
This work describes an automatic query refinement technique, which focuses on improving precision of the top ranked documents. The terms used for refinement are lexical affinities (LAs), pairs of closely related words which contain exactly one of the original query terms. Adding these terms to the query is equivalent to re-ranking search results, thus, precision is improved while recall is preserved. We describe a novel method that selects the most "informative" LAs for refinement, namely, those LAs that best separate relevant documents from irrelevant documents in the set of results. The information gain of candidate LAs is determined using unsupervised estimation that is based on the scoring function of the search engine. This method is thus fully automatic and its quality depends on the quality of the scoring function. Experiments we conducted with TREC data clearly show a significant improvement in the precision of the top ranked documents.