Detecting Fake News Over Online Social Media via Domain Reputations and Content Understanding

Detecting Fake News Over Online Social Media via Domain Reputations and Content Understanding
复制标题

DOI:
10.26599/tst.2018.9010139
复制
发表时间:
2020-02-01
影响因子:
6.6
通讯作者:
Yang, Bo
Yang, Bo
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xu, Kuai;Wang, Feng;Yang, Bo

文献摘要

被引文献

相似文献

最近,假新闻利用网络社交媒体的力量和规模,有效地传播错误信息,不仅侵蚀了人们对传统媒体和新闻的信任,而且还操纵了公众的观点和情绪。由于真假新闻的细微差别,检测假新闻是一项艰巨的挑战。作为打击假新闻的第一步,本文从两个角度描述了数百个流行的假新闻和真实新闻,这些新闻通过Facebook上的份额、反应和评论来衡量:领域声誉和内容理解。我们的域名声誉分析表明,假新闻和真新闻发布者的网站在注册行为、注册时间、域名排名和域名受欢迎程度等方面表现出不同的特征。此外,假新闻往往会在一段时间后从网络上消失。假新闻和真实新闻语料库的内容表征表明,简单地应用词频逆文档频率(tf-idf)和潜在狄利let分配(LDA)主题建模在假新闻检测中是低效的,而利用词和词向量探索文档相似度是预测假新闻和真实新闻的一个很有前途的方向。据我们所知,这是第一次系统地研究假新闻和真实新闻的领域声誉和内容特征,这将为有效检测社交媒体上的假新闻提供关键见解。
Fake news has recently leveraged the power and scale of online social media to effectively spread misinformation which not only erodes the trust of people on traditional presses and journalisms, but also manipulates the opinions and sentiments of the public. Detecting fake news is a daunting challenge due to subtle difference between real and fake news. As a first step of fighting with fake news, this paper characterizes hundreds of popular fake and real news measured by shares, reactions, and comments on Facebook from two perspectives: domain reputations and content understanding. Our domain reputation analysis reveals that the Web sites of the fake and real news publishers exhibit diverse registration behaviors, registration timing, domain rankings, and domain popularity. In addition, fake news tends to disappear from the Web after a certain amount of time. The content characterizations on the fake and real news corpus suggest that simply applying term frequency-inverse document frequency (tf-idf) and Latent Dirichlet Allocation (LDA) topic modeling is inefficient in detecting fake news, while exploring document similarity with the term and word vectors is a very promising direction for predicting fake and real news. To the best of our knowledge, this is the first effort to systematically study domain reputations and content characteristics of fake and real news, which will provide key insights for effectively detecting fake news on social media.