Behind the cues: A benchmarking study for fake news detection

Behind the cues: A benchmarking study for fake news detection
复制标题

DOI:
10.1016/j.eswa.2019.03.036
复制
发表时间:
2019-08-15
影响因子:
8.5
通讯作者:
Karadais, Panagiotis
Karadais, Panagiotis
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gravanis, Georgios;Vakali, Athena;Karadais, Panagiotis

文献摘要

被引文献

相似文献

假新闻已经成为信息社会中一个影响巨大的问题,因为造假者的内容传播持续而激烈。新闻源中的信息质量真实性值得怀疑,需要自动化工具来检测虚假新闻文章。由于伪造者的面孔很多,创建这样的工具是一个具有挑战性的问题。在这项工作中,我们提出了一个使用基于内容的特征和机器学习(ML)算法的假新闻检测模型。为了在最准确的模型中得出结论,我们还评估了用于欺骗检测和词嵌入的几个特征集。此外,我们测试了最流行的ML分类器,并研究了在AdaBoost和Bagging等集成ML方法下可能实现的改进。一组广泛的早期数据源已被用于特征集和ML分类器的实验和评估。此外,我们引入了一个新的文本语料库,“无偏”(UNB)数据集,它集成了各种新闻来源,并满足一些标准和规则,以避免偏见的结果在分类任务。我们的实验结果表明,使用增强的语言特征集与词嵌入沿着与集成算法和支持向量机(SVM)能够分类假新闻具有高准确性。(C)2019爱思唯尔有限公司版权所有。
Fake news has become a problem of great impact in our information driven society because of the continuous and intense fakesters content distribution. Information quality in news feeds is under questionable veracity calling for automated tools to detect fake news articles. Due to many faces of fakesters, creating such tool is a challenging problem. In this work, we propose a model for fake news detection using content based features and Machine Learning (ML) algorithms. To conclude in most accurate model we evaluate several feature sets proposed for deception detection and word embeddings as well. Moreover, we test the most popular ML classifiers and investigate the possible improvement reached under ensemble ML methods such as AdaBoost and Bagging. An extensive set of earlier data sources has been used for experimentation and evaluation of both feature sets and ML classifiers. Moreover, we introduce a new text corpus, the "UNBiased" (UNB) dataset, which integrates various news sources and fulfills several standards and rules to avoid biased results in classification task. Our experimental results show that the use of an enhanced linguistic feature set with word embeddings along with ensemble algorithms and Support Vector Machines (SVMs) is capable to classify fake news with high accuracy. (C) 2019 Elsevier Ltd. All rights reserved.