An apple-to-apple comparison of Learning-to-rank algorithms in terms of Normalized Discounted Cumulative Gain

An apple-to-apple comparison of Learning-to-rank algorithms in terms of Normalized Discounted Cumulative Gain
复制标题

DOI:
--
复制
发表时间:
2012-08
期刊:
--
影响因子:
--
通讯作者:
R. Busa-Fekete;Gy¨orgy Szarvas;Tam´as ´Eltet˝o;B. K´egl
R. Busa-Fekete;Gy¨orgy Szarvas;Tam´as ´Eltet˝o;B. K´egl
中科院分区:
其他
文献类型:
--
作者:
R. Busa-Fekete;Gy¨orgy Szarvas;Tam´as ´Eltet˝o;B. K´egl

文献摘要

被引文献

相似文献

归一化贴现累积增益(NDCG)是一种广泛应用于LTR系统的评价指标。NDCG设计用于对具有多个相关级别的任务进行排序。有许多免费的开源工具可用于计算排名结果列表的NDCG分数。尽管NDCG的定义是明确的,但不同的工具会对具有某些属性的排名列表产生不同的分数,从而恶化了许多已发表论文的实证检验,从而使不同研究发表的实证结果的比较难以比较。在本研究中,首先,我们确定了各种公开可用的NDCG评估工具之间的主要差异。其次,基于一组使用LTR研究中常见基准数据集和6种不同LTR算法的比较实验,我们展示了这些差异如何影响不同算法的整体性能以及用于比较不同系统的最终分数。
The Normalized Discounted Cumulative Gain (NDCG) is a widely used evaluation metric for learning-to-rank (LTR) systems. NDCG is designed for ranking tasks with more than one relevance levels. There are many freely available, open source tools for computing the NDCG score for a ranked result list. Even though the definition of NDCG is unambiguous, the various tools can produce different scores for ranked lists with certain properties, deteriorating the empirical tests in many published papers and thereby making the comparison of empirical results published in different studies difficult to compare. In this study, first, we identify the major differences between the various publicly available NDCG evaluation tools. Second, based on a set of comparative experiments using a common benchmark dataset in LTR research and 6 different LTR algorithms, we demonstrate how these differences affect the overall performance of different algorithms and the final scores that are used to compare different systems.