Why batch and user evaluations do not give the same results
Why batch and user evaluations do not give the same results
复制标题
为什么批次评估和用户评估不会给出相同的结果
DOI:
10.1145/383952.383992
复制
发表时间:
2001
期刊:
影响因子:
--
通讯作者:
W. Hersh
中科院分区:
文献类型:
--
作者:
A. Turpin;W. Hersh
Much system-oriented evaluation of information retrieval systems has used the Cranfield approach based upon queries run against test collections in a batch mode. Some researchers have questioned whether this approach can be applied to the real world, but little data exists for or against that assertion. We have studied this question in the context of the TREC Interactive Track. Previous results demonstrated that improved performance as measured by relevance-based metrics in batch studies did not correspond with the results of outcomes based on real user searching tasks. The experiments in this paper analyzed those results to determine why this occurred. Our assessment showed that while the queries entered by real users into systems yielding better results in batch studies gave comparable gains in ranking of relevant documents for those users, they did not translate into better performance on specific tasks. This was most likely due to users being able to adequately find and utilize relevant documents ranked further down the output list.