Why batch and user evaluations do not give the same results

Why batch and user evaluations do not give the same results
复制标题

为什么批次评估和用户评估不会给出相同的结果

DOI:
10.1145/383952.383992
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
W. Hersh
W. Hersh
中科院分区:
--
文献类型:
--
作者:
A. Turpin;W. Hersh

文献摘要

被引文献

相似文献

许多面向系统的信息检索系统评估使用了克兰菲尔德方法,该方法基于以批处理模式对测试集合运行的查询。一些研究人员质疑这种方法是否可以应用于现实世界,但支持或反对这一断言的数据很少。我们在TREC互动轨道的背景下研究了这个问题。先前的结果表明,批量研究中基于相关性的指标所衡量的性能改善与基于真实用户搜索任务的结果不一致。本文中的实验分析了这些结果,以确定为什么会发生这种情况。我们的评估表明,尽管真实用户输入的查询在批量研究中产生更好的结果,在这些用户的相关文件排名方面有类似的收益,但它们并没有转化为在具体任务上的更好表现。这很可能是因为用户能够充分查找和利用在产出列表中排名靠后的相关文件。
Much system-oriented evaluation of information retrieval systems has used the Cranfield approach based upon queries run against test collections in a batch mode. Some researchers have questioned whether this approach can be applied to the real world, but little data exists for or against that assertion. We have studied this question in the context of the TREC Interactive Track. Previous results demonstrated that improved performance as measured by relevance-based metrics in batch studies did not correspond with the results of outcomes based on real user searching tasks. The experiments in this paper analyzed those results to determine why this occurred. Our assessment showed that while the queries entered by real users into systems yielding better results in batch studies gave comparable gains in ranking of relevant documents for those users, they did not translate into better performance on specific tasks. This was most likely due to users being able to adequately find and utilize relevant documents ranked further down the output list.