FA-KES: A Fake News Dataset around the Syrian War

FA-KES: A Fake News Dataset around the Syrian War
复制标题

FA-KES:围绕叙利亚战争的假新闻数据集

DOI:
10.1609/icwsm.v13i01.3254
复制
发表时间:
2019
影响因子:
3.2
通讯作者:
May Farah
May Farah
中科院分区:
法学4区
文献类型:
--
作者:
Fatima K. Abu Salem;Roaa Al Feel;Shady Elbassuoni;Mohamad Jaber;May Farah

文献摘要

被引文献

相似文献

目前大多数可用的假新闻数据集都围绕美国政治、夹带新闻或讽刺。它们通常是从事实核查网站上抓取的,这些文章由人类专家标记。在本文中,我们提出了 FA-KES,一个围绕叙利亚战争的假新闻数据集。考虑到战争事件新闻报道的特殊性质,以及缺乏可从中抓取手动标记的新闻文章的可用来源,我们认为专门为该领域构建的假新闻数据集至关重要。为了确保数据集涵盖叙利亚战争的各个方面,我们的数据集包含来自代表动员媒体、忠诚媒体和各种印刷媒体的多家媒体的新闻文章。为了避免手动将新闻文章标记为真或假的困难且通常是主观的任务,我们采用半监督事实检查方法来标记数据集中的新闻文章。在众包的帮助下,人类贡献者被提示提取具体且易于提取的信息,这些信息有助于将给定的文章与从叙利亚违法行为文档中心获得的代表“真实情况”的信息进行匹配。然后,使用无监督机器学习将提取的信息聚类为两个独立的集合。结果是一个经过仔细注释的数据集,其中包含 804 篇标记为真或假的文章,非常适合训练机器学习模型来预测新闻文章的可信度。我们的数据集可在 https://doi.org/10.5281/zenodo.2607278 上公开获取。尽管我们的数据集主要关注叙利亚危机,但它可以用于训练机器学习模型以检测其他相关领域的假新闻。此外,我们用于获取数据集的框架足够通用,可以用于围绕军事冲突构建其他假新闻数据集,前提是有一些相应的地面事实可用。
Most currently available fake news datasets revolve around US politics, entrainment news or satire. They are typically scraped from fact-checking websites, where the articles are labeled by human experts. In this paper, we present FA-KES, a fake news dataset around the Syrian war. Given the specific nature of news reporting on incidents of wars and the lack of available sources from which manually-labeled news articles can be scraped, we believe a fake news dataset specifically constructed for this domain is crucial. To ensure a balanced dataset that covers the many facets of the Syrian war, our dataset consists of news articles from several media outlets representing mobilisation press, loyalist press, and diverse print media. To avoid the difficult and often-subjective task of manually labeling news articles as true or fake, we employ a semi-supervised fact-checking approach to label the news articles in our dataset. With the help of crowdsourcing, human contributors are prompted to extract specific and easy-to-extract information that helps match a given article to information representing “ground truth” obtained from the Syrian Violations Documentation Center. The information extracted is then used to cluster the articles into two separate sets using unsupervised machine learning. The result is a carefully annotated dataset consisting of 804 articles labeled as true or fake and that is ideal for training machine learning models to predict the credibility of news articles. Our dataset is publicly available at https://doi.org/10.5281/zenodo.2607278. Although our dataset is focused on the Syrian crisis, it can be used to train machine learning models to detect fake news in other related domains. Moreover, the framework we used to obtain the dataset is general enough to be used to build other fake news datasets around military conflicts, provided there is some corresponding ground-truth available.