Regularizing ad hoc retrieval scores
Regularizing ad hoc retrieval scores
复制标题
DOI:
10.1145/1099554.1099722
复制
发表时间:
2005-10
期刊:
影响因子:
--
通讯作者:
Fernando Diaz
中科院分区:
文献类型:
--
作者:
Fernando Diaz
The cluster hypothesis states: closely related documents tend to be relevant to the same request. We exploit this hypothesis directly by adjusting ad hoc retrieval scores from an initial retrieval so that topically related documents receive similar scores. We refer to this process as score regularization. Score regularization can be presented as an optimization problem, allowing the use of results from semi-supervised learning. We demonstrate that regularized scores consistently and significantly rank documents better than unregularized scores, given a variety of initial retrieval algorithms. We evaluate our method on two large corpora across a substantial number of topics.