Detecting opinion spams and fake news using text classification

Detecting opinion spams and fake news using text classification
复制标题

DOI:
10.1002/spy2.9
复制
发表时间:
2018-01-01
影响因子:
1.9
通讯作者:
Saad, Sherif
Saad, Sherif
中科院分区:
其他
文献类型:
--
作者:
Ahmed, Hadeer;Traore, Issa;Saad, Sherif

文献摘要

被引文献

相似文献

近年来,假新闻和假评论等欺骗性内容,也被称为意见垃圾,越来越成为网络用户的危险前景。虚假评论影响了消费者和商店。此外,假新闻问题在2016年引起了人们的关注,特别是在上一次美国总统选举之后。假评论和假新闻是一种密切相关的现象,因为两者都包括撰写和传播虚假信息或信念。意见垃圾邮件问题在几年前首次提出,但由于用户生成内容的丰富性,它已迅速成为一个不断增长的研究领域。现在任何人都可以很容易地在网上写虚假评论或写假新闻。最大的挑战是缺乏一种有效的方法来区分真实的评论和虚假评论;即使是人类也往往无法区分。在本文中,我们引入了一个新的n-gram模型来自动检测虚假内容,特别关注虚假评论和假新闻。我们研究和比较了2种不同的特征提取技术和6种机器学习分类技术。使用现有的公共数据集和新引入的假新闻数据集进行的实验评估表明,与最先进的方法相比,性能非常令人鼓舞和改进。
In recent years, deceptive content such as fake news and fake reviews, also known as opinion spams, have increasingly become a dangerous prospect for online users. Fake reviews have affected consumers and stores alike. Furthermore, the problem of fake news has gained attention in 2016, especially in the aftermath of the last U.S. presidential elections. Fake reviews and fake news are a closely related phenomenon as both consist of writing and spreading false information or beliefs. The opinion spam problem was formulated for the first time a few years ago, but it has quickly become a growing research area due to the abundance of user-generated content. It is now easy for anyone to either write fake reviews or write fake news on the web. The biggest challenge is the lack of an efficient way to tell the difference between a real review and a fake one; even humans are often unable to tell the difference. In this paper, we introduce a new n-gram model to detect automatically fake contents with a particular focus on fake reviews and fake news. We study and compare 2 different features extraction techniques and 6 machine learning classification techniques. Experimental evaluation using existing public datasets and a newly introduced fake news dataset indicate very encouraging and improved performances compared to the state-of-the-art methods.