Unanimity-Aware Gain for Highly Subjective Assessments

Unanimity-Aware Gain for Highly Subjective Assessments
复制标题

高度主观评估的一致感知收益

DOI:
--
复制
发表时间:
2017
期刊:
EVIA@NTCIR
影响因子:
--
通讯作者:
T. Sakai
T. Sakai
中科院分区:
--
文献类型:
--
作者:
T. Sakai

文献摘要

被引文献

相似文献

IR任务多样化:对诸如社交媒体帖子之类的项目的人类评估可能是高度主观的,在这种情况下,有必要为每个项目雇用许多评估者以反映他们的不同观点。例如,对于给定目的,推文的价值可以由(比如说)十个评估者来判断,并且他们的评级可以被求和以确定其增益值,用于计算分级相关性评估度量。在本研究中,我们提出了这种方法的一个简单的变体,它考虑到这样一个事实,即一些项目获得一致的评级,而另一些则更具争议性。我们基于真实的基于社交媒体的IR任务数据生成模拟评级,以检查我们的安全性感知方法对系统排名和统计显著性的影响。我们的研究结果表明,即使当它对增益值的影响保持在最小值时,合并的重复性也可以影响统计显著性测试结果。此外,由于我们的模拟评分没有考虑评估员实际评分中存在的相关性,我们的实验可能低估了将相似性引入评估的效果。因此,如果研究人员接受一致投票应该比有争议的投票更有价值,那么我们提出的方法可能值得采用。
IR tasks have diversied: human assessments of items such as social media posts can be highly subjective, in which case it becomes necessary to hire many assessors per item to reect their diverse views. For example, the value of a tweet for a given purpose may be judged by (say) ten assessors, and their ratings could be summed up to dene its gain value for computing a graded-relevance evaluation measure. In the present study, we propose a simple variant of this approach, which takes into account the fact that some items receive unanimous ratings while others are more controversial. We generate simulated ratings based on a real social-media-based IR task data to examine the eect of our unanimity-aware approach on the system ranking and on statistical signicance. Our results show that incorporating unanimity can aect statistical signicance test results even when its impact on the gain value is kept to a minimum. Moreover, since our simulated ratings do not consider the correlation present in the assessors’ actual ratings, our experiments probably underestimate the eect of introducing unanimity into evaluation. Hence, if researchers accept that unanimous votes should be valued more highly than controversial ones, then our proposed approach may be worth incorporating.