Good Evaluation Measures based on Document Preferences

Good Evaluation Measures based on Document Preferences
复制标题

DOI:
10.1145/3397271.3401115
复制
发表时间:
2020-07
期刊:
Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval
影响因子:
--
通讯作者:
T. Sakai;Zhaohao Zeng
T. Sakai;Zhaohao Zeng
中科院分区:
其他
文献类型:
--
作者:
T. Sakai;Zhaohao Zeng

文献摘要

被引文献

相似文献

对于信息检索系统的离线评估,一些研究人员建议使用成对的文件偏好评估,而不是单个文件的相关性评估,因为这样可能更容易让评价者做出相对决定,而不是绝对决定。已经提出了基于偏好的简单评估措施,如Ppref和Wpref,但在过去十年中,这种措施没有得到任何广泛使用。其中一个原因可能是,尽管据报道,这些新的衡量标准与基于绝对评估的传统衡量标准或多或少类似,但它们是否真的与用户对搜索引擎结果页面(SERP)的感知一致尚不清楚。在正式定义了两类基于偏好的度量后,本研究就解决了这个问题,这两类度量被称为Pref度量和Δ度量。我们表明,在与用户的SERP偏好的一致性方面,这些衡量标准中最好的至少与一般评估者一样好,并且隐性文档偏好(即,由检索一个文档但不检索另一个文档的SERP建议的文档偏好)比显性偏好(即,由检索一个文档高于另一个文档的SERP建议的文档偏好)发挥着更重要的作用。我们已经发布了包含119,646个文档偏好的数据集,以便IR社区可以进一步探索基于文档偏好的评估的可行性。
For offline evaluation of IR systems, some researchers have proposed to utilise pairwise document preference assessments instead of relevance assessments of individual documents, as it may be easier for assessors to make relative decisions rather than absolute ones. Simple preference-based evaluation measures such as ppref and wpref have been proposed, but the past decade did not see any wide use of such measures. One reason for this may be that, while these new measures have been reported to behave more or less similarly to traditional measures based on absolute assessments, whether they actually align with the users' perception of search engine result pages (SERPs) has been unknown. The present study addresses exactly this question, after formally defining two classes of preference-based measures called Pref measures and Δ-measures. We show that the best of these measures perform at least as well as an average assessor in terms of agreement with users' SERP preferences, and that implicit document preferences (i.e., those suggested by a SERP that retrieves one document but not the other) play a much more important role than explicit preferences (i.e., those suggested by a SERP that retrieves one document above the other). We have released our data set containing 119,646 document preferences, so that the feasibility of document preferenced-based evaluation can be further pursued by the IR community.