Using hit curves to compare search algorithm performance.

Using hit curves to compare search algorithm performance.
复制标题

使用命中曲线来比较搜索算法的性能。

DOI:
10.1016/j.jbi.2005.12.007
复制
发表时间:
2007
影响因子:
4.5
通讯作者:
Bernstam,ElmerV
Bernstam,ElmerV
中科院分区:
医学3区
文献类型:
--
作者:
Herskovic,JorgeR;Iyengar,MSriram;Bernstam,ElmerV

文献摘要

相似文献

数据库继续增长,但可用于评估信息检索系统的指标并没有改变。大型的数据库,如MEDLINE和万维网,包含许多相关的文档,可用于常见的查询。因此,排名越来越重要,并且成功的信息检索系统(诸如Google)已经强调了排名。然而,现有的评估指标,如精度和召回率,不直接考虑排名。本文介绍了一种新的方法来衡量信息检索性能加权命中曲线适用于统计检测领域,以反映多个理想的特性,如相关性,重要性和方法质量。在统计检测中,已经提出命中曲线来表示检测过程期间感兴趣事件的发生。类似地,命中曲线可用于研究相关文档在大型结果集中的位置。我们描述命中曲线的信息检索的正式模型,显示命中曲线如何表示系统的性能,包括排名,并定义方法来统计比较多个系统的性能使用命中曲线。我们提供了传统指标不如命中曲线适合的示例场景,并得出结论,命中曲线可能有助于评估从排名性能至关重要的大型集合中的检索。
Databases continue to grow but the metrics available to evaluate information retrieval systems have not changed. Large collections such as MEDLINE and the World Wide Web contain many relevant documents for common queries. Ranking is therefore increasingly important and successful information retrieval systems, such as Google, have emphasized ranking. However, existing evaluation metrics such as precision and recall, do not directly account for ranking. This paper describes a novel way of measuring information retrieval performance using weighted hit curves adapted from the field of statistical detection to reflect multiple desirable characteristics such as relevance, importance, and methodologic quality. In statistical detection, hit curves have been proposed to represent occurrence of interesting events during a detection process. Similarly, hit curves can be used to study the position of relevant documents within large result sets. We describe hit curves in light of a formal model of information retrieval, show how hit curves represent system performance including ranking, and define ways to statistically compare performance of multiple systems using hit curves. We provide example scenarios where traditional measures are less suitable than hit curves and conclude that hit curves may be useful for evaluating retrieval from large collections where ranking performance is crucial.