Adapting boosting for information retrieval measures

Adapting boosting for information retrieval measures
复制标题

DOI:
10.1007/s10791-009-9112-1
复制
发表时间:
2010-06
期刊:
Information Retrieval
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

被引文献

相似文献

我们提出了一种新的排名算法,结合了以前的两种方法的优势:提升树分类,和LambdaRank,这已被证明是经验最佳的广泛使用的信息检索措施。我们的算法基于增强回归树,尽管这些想法适用于任何弱学习者,并且在训练和测试阶段都比最先进的算法快得多,精度相当。我们还展示了如何找到任何两个排名的最佳线性组合,我们使用这种方法来解决线搜索问题,在提升。此外,我们表明,从以前训练过的模型开始,并使用其残差进行提升,这是一种有效的模型自适应技术,并且我们对网络搜索中的一个特别紧迫的问题给出了显着改进的结果-训练排名市场只有少量的标记数据可用,给定一个来自更大市场的更多数据训练的排名。
We present a new ranking algorithm that combines the strengths of two previous methods: boosted tree classification, and LambdaRank, which has been shown to be empirically optimal for a widely used information retrieval measure. Our algorithm is based on boosted regression trees, although the ideas apply to any weak learners, and it is significantly faster in both train and test phases than the state of the art, for comparable accuracy. We also show how to find the optimal linear combination for any two rankers, and we use this method to solve the line search problem exactly during boosting. In addition, we show that starting with a previously trained model, and boosting using its residuals, furnishes an effective technique for model adaptation, and we give significantly improved results for a particularly pressing problem in web search—training rankers for markets for which only small amounts of labeled data are available, given a ranker trained on much more data from a larger market.