The effect of threshold priming and need for cognition on relevance calibration and assessment

The effect of threshold priming and need for cognition on relevance calibration and assessment
复制标题

阈值启动和认知需求对相关性校准和评估的影响

DOI:
10.1145/2484028.2484090
复制
发表时间:
2013
期刊:
Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval
影响因子:
--
通讯作者:
William Webber
William Webber
中科院分区:
--
文献类型:
--
作者:
Falk Scholer;D. Kelly;Wan;H. Lee;William Webber

文献摘要

被引文献

相似文献

需要对文档相关性进行人工评估,以构建测试集合、进行特别评估和培训文本分类器。然而,以不同的顺序向评审员展示文档可能会导致不同的评估结果。我们考察了定义术语{阈值启动},看到不同程度的相关文件,对人们的相关性校准的影响。与会者对包含高度相关、中等相关或不相关文件的文件序言的相关性进行了判断,随后对混合相关性的文件进行了共同结语。我们观察到,只接触前言中不相关文件的参与者对序言和结语文件的平均相关性得分显著高于接触序言中相关性中等或高度相关文件的参与者。我们还考察了DefineTerm(认知需要)是如何影响相关性评估的,它是衡量一个人参与努力认知活动程度的个体差异衡量标准。高认知需求参与者与专家评审员的一致性显著高于低认知需求参与者。我们的发现表明,评审员应该在评判过程的早期接触到来自多个相关性水平的文件,以便以平衡的方式校准他们的相关性阈值,而个体差异测量可能是筛选评审员的有用方式。
Human assessments of document relevance are needed for the construction of test collections, for ad-hoc evaluation, and for training text classifiers. Showing documents to assessors in different orderings, however, may lead to different assessment outcomes. We examine the effect that \defineterm{threshold priming}, seeing varying degrees of relevant documents, has on people's calibration of relevance. Participants judged the relevance of a prologue of documents containing highly relevant, moderately relevant, or non-relevant ocuments, followed by a common epilogue of documents of mixed relevance. We observe that participants exposed to only non-relevant documents in the prologue assigned significantly higher average relevance scores to prologue and epilogue documents than participants exposed to moderately or highly relevant documents in the prologue. We also examine how \defineterm{need for cognition}, an individual difference measure of the extent to which a person enjoys engaging in effortful cognitive activity, impacts relevance assessments. High need for cognition participants had a significantly higher level of agreement with expert assessors than low need for cognition participants did. Our findings indicate that assessors should be exposed to documents from multiple relevance levels early in the judging process, in order to calibrate their relevance thresholds in a balanced way, and that individual difference measures might be a useful way to screen assessors.