Regularizing ad hoc retrieval scores

Regularizing ad hoc retrieval scores
复制标题

DOI:
10.1145/1099554.1099722
复制
发表时间:
2005-10
期刊:
--
影响因子:
--
通讯作者:
Fernando Diaz
Fernando Diaz
中科院分区:
其他
文献类型:
--
作者:
Fernando Diaz

文献摘要

被引文献

相似文献

集群假设指出:密切相关的文档往往与相同的请求相关。我们直接利用这一假设,通过调整初始检索的特别检索分数,使主题相关的文档获得类似的分数。我们将这一过程称为分数正规化。分数正则化可以表示为一个优化问题,允许使用半监督学习的结果。我们证明,在给定各种初始检索算法的情况下,规则分数一致且显著地比非规则分数更好地对文档进行排名。我们在两个大型语料库上对我们的方法进行了评估,涉及大量主题。
The cluster hypothesis states: closely related documents tend to be relevant to the same request. We exploit this hypothesis directly by adjusting ad hoc retrieval scores from an initial retrieval so that topically related documents receive similar scores. We refer to this process as score regularization. Score regularization can be presented as an optimization problem, allowing the use of results from semi-supervised learning. We demonstrate that regularized scores consistently and significantly rank documents better than unregularized scores, given a variety of initial retrieval algorithms. We evaluate our method on two large corpora across a substantial number of topics.