Investigating Differences in Crowdsourced News Credibility Assessment: Raters, Tasks, and Expert Criteria

Investigating Differences in Crowdsourced News Credibility Assessment: Raters, Tasks, and Expert Criteria
复制标题

调查众包新闻可信度评估的差异:评估者、任务和专家标准

DOI:
10.1145/3415164
复制
发表时间:
2020
影响因子:
--
通讯作者:
Mitra, Tanushree
Mitra, Tanushree
中科院分区:
--
文献类型:
--
作者:
Bhuiyan, Md Momen;Zhang, Amy X.;Sehat, Connie Moon;Mitra, Tanushree

文献摘要

相似文献

有关气候变化和疫苗安全等关键问题的错误信息经常在在线社交和搜索平台上被放大。外行人对内容可信度评估的众包已被提议作为一种通过尝试大规模复制专家的评估来打击错误信息的策略。在这项工作中,我们调查了大众与专家对新闻可信度的评估,以了解他们之间的评级何时以及如何存在差异。我们收集了 4,000 多个可信度评估的数据集,这些评估来自 2 个人群群体(新闻系学生和 Upwork 工作者)以及 2 个专家组(记者和科学家),涉及气候科学相关的 50 篇新闻文章,而气候科学是一个公众舆论与专家共识之间普遍脱节的话题。通过检查评级,我们发现由于人群的构成(例如评级者的人口统计数据和政治倾向)以及人群被分配评级的任务范围(例如文章的类型和出版物的党派倾向)而导致表现的差异。最后,我们发现由于新闻专家与科学专家使用的专家标准不同而导致专家评估之间的差异,这些差异可能会导致人群差异,但这也提出了一种通过设计适合特定专家标准的人群任务来缩小差距的方法。根据这些发现,我们概述了未来的研究方向,以更好地设计针对特定人群和内容类型的人群流程。
Misinformation about critical issues such as climate change and vaccine safety is oftentimes amplified on online social and search platforms. The crowdsourcing of content credibility assessment by laypeople has been proposed as one strategy to combat misinformation by attempting to replicate the assessments of experts at scale. In this work, we investigate news credibility assessments by crowds versus experts to understand when and how ratings between them differ. We gather a dataset of over 4,000 credibility assessments taken from 2 crowd groups---journalism students and Upwork workers---as well as 2 expert groups---journalists and scientists---on a varied set of 50 news articles related to climate science, a topic with widespread disconnect between public opinion and expert consensus. Examining the ratings, we find differences in performance due to the makeup of the crowd, such as rater demographics and political leaning, as well as the scope of the tasks that the crowd is assigned to rate, such as the genre of the article and partisanship of the publication. Finally, we find differences between expert assessments due to differing expert criteria that journalism versus science experts use---differences that may contribute to crowd discrepancies, but that also suggest a way to reduce the gap by designing crowd tasks tailored to specific expert criteria. From these findings, we outline future research directions to better design crowd processes that are tailored to specific crowds and types of content.