To Check or Not to Check: Syntax, Semantics, and Context in the Language of Check-Worthy Claims

To Check or Not to Check: Syntax, Semantics, and Context in the Language of Check-Worthy Claims
复制标题

DOI:
10.1007/978-3-030-28577-7_23
复制
发表时间:
2019-09
期刊:
--
影响因子:
--
通讯作者:
Chaoyuan Zuo;A. Karakas;Ritwik Banerjee
Chaoyuan Zuo;A. Karakas;Ritwik Banerjee
中科院分区:
其他
文献类型:
--
作者:
Chaoyuan Zuo;A. Karakas;Ritwik Banerjee

文献摘要

被引文献

相似文献

由于社交媒体的普遍使用,信息的传播得到了令人信服的推动,错误信息的传播也是如此。庞大的数据量使得专家驱动的人工事实核查的传统方法在很大程度上是不可行的。因此,近年来对计算语言学和数据驱动算法进行了探索。尽管取得了这一进展,但确定需要检查的内容并确定其优先顺序的工作几乎没有受到关注。鉴于专家驱动的人工干预可能仍然是事实核查的一个重要组成部分,特别是在特定领域(如政治、环境科学),这种确定和优先次序至关重要。对“值得核查的”索赔进行成功的算法排序可以帮助专家进行环路事实核查系统,从而减少专家的工作量,同时仍能处理最突出的错误信息。在这项工作中,我们探索了语言句法、语义和单词的上下文意义如何在决定索赔的可检验性方面发挥作用。在CLEF-2018事实核查实验室的可验证性任务中,我们的初步实验使用了明确的文体特征和在英语语言数据集上的简单单词嵌入,其中我们的主要解决方案在平均平均精度、R精度、倒数排名和多值精度atk方面优于其他系统。在这里,我们提出了这种方法的扩展,具有更复杂的单词嵌入,并报告了在这项任务中的进一步改进。
As the spread of information has received a compelling boost due to pervasive use of social media, so has the spread of misinformation. The sheer volume of data has rendered the traditional methods of expert-driven manual fact-checking largely infeasible. As a result, computational linguistics and data-driven algorithms have been explored in recent years. Despite this progress, identifying and prioritizingwhatneeds to be checked has received little attention. Given that expert-driven manual intervention is likely to remain an important component of fact-checking, especially in specific domains (e.g., politics, environmental science), this identification and prioritization is critical. A successful algorithmic ranking of “check-worthy” claims can help an expert-in-the-loop fact-checking system, thereby reducing the expert’s workload while still tackling the most salient bits of misinformation. In this work, we explore how linguistic syntax, semantics, and the contextual meaning of words play a role in determining the check-worthiness of claims. Our preliminary experiments used explicit stylometric features and simple word embeddings on the English language dataset in the Check-worthiness task of the CLEF-2018 Fact-Checking Lab, where our primary solution outperformed the other systems in terms of the mean average precision,R-precision, reciprocal rank, and precision atkfor multiple valuesk. Here, we present an extension of this approach with more sophisticated word embeddings and report further improvements in this task.