LETOR: Benchmark Dataset for Research on Learning to Rank for Information Retrieval

LETOR: Benchmark Dataset for Research on Learning to Rank for Information Retrieval
复制标题

DOI:
--
复制
发表时间:
2007
影响因子:
--
通讯作者:
Tie-Yan Liu;Jun Xu;Tao Qin;Wen-Ying Xiong;Hang Li
Tie-Yan Liu;Jun Xu;Tao Qin;Wen-Ying Xiong;Hang Li
中科院分区:
--
文献类型:
--
作者:
Tie-Yan Liu;Jun Xu;Tao Qin;Wen-Ying Xiong;Hang Li

文献摘要

被引文献

相似文献

本文研究了信息检索中的排序学习问题。排序是信息检索的核心问题,采用机器学习技术来学习排序函数被认为是一种很有前途的方法。不幸的是,没有基准数据集,可以用来比较现有的学习算法和评估新提出的算法,这阻碍了相关的研究。为了解决这个问题,我们构建了一个称为LETOR的基准数据集,并将其分发给研究社区。具体来说,我们已经从现有的数据集广泛使用的IR,即OHSUMED和TREC数据的LETOR数据。这两个集合包含查询、检索到的文档的内容以及关于文档与查询的相关性的人类判断。我们已经从数据集中提取了特征,包括常规特征,如词频,逆文档频率,BM 25和IR的语言模型,以及SIGIR最近提出的特征,如HostRank,特征传播和主题PageRank。然后,我们将LETOR与提取的特征、查询和相关性判断打包。我们还提供了几种最先进的学习结果,以根据数据对算法进行排名。本文详细介绍了LETOR。
This paper is concerned with learning to rank for information retrieval (IR). Ranking is the central problem for information retrieval, and employing machine learning techniques to learn the ranking function is viewed as a promising approach to IR. Unfortunately, there was no benchmark dataset that could be used in comparison of existing learning algorithms and in evaluation of newly proposed algorithms, which stood in the way of the related research. To deal with the problem, we have constructed a benchmark dataset referred to as LETOR and distributed it to the research communities. Specifically we have derived the LETOR data from the existing data sets widely used in IR, namely, OHSUMED and TREC data. The two collections contain queries, the contents of the retrieved documents, and human judgments on the relevance of the documents with respect to the queries. We have extracted features from the datasets, including both conventional features, such as term frequency, inverse document frequency, BM25, and language models for IR, and features proposed recently at SIGIR, such as HostRank, feature propagation, and topical PageRank. We have then packaged LETOR with the extracted features, queries, and relevance judgments. We have also provided the results of several state-ofthe-arts learning to rank algorithms on the data. This paper describes in details about LETOR.