Development of a 'fake news' machine learning classifier and a dataset for its testing

Development of a 'fake news' machine learning classifier and a dataset for its testing
复制标题

开发“假新闻”机器学习分类器及其测试数据集

DOI:
10.1117/12.2520131
复制
发表时间:
2019
期刊:
Proceedings of SPIE
影响因子:
--
通讯作者:
Straub, Jeremy
Straub, Jeremy
中科院分区:
--
文献类型:
--
作者:
Fleck, William;Snell, Nicholas;Traylor, Terry;Straub, Jeremy

文献摘要

被引文献

相似文献

在2016年美国总统大选之后,包含虚假信息但被呈现为事实准确的捏造新闻故事(通常称为“假新闻”)引起了极大的兴趣和媒体的关注。虽然选举期间发生的事情的全部细节尚不清楚,但似乎有多个团体利用社交媒体传播虚假信息,这些信息被包装在捏造的新闻文章中,被认为是真实的。一些人认为,这场运动对选举产生了重大影响。此外,2016年美国总统大选远非唯一一场假新闻发挥明显作用的竞选活动。在本文中,工作的反假新闻的研究工作。从长远来看,该项目的重点是建立一个潜在的欺骗性虚假内容的指示和警告系统。作为该项目的一部分,人工分类的合法和欺骗性新闻文章的数据集被策展。识别和讨论了人工分类项目确定的合法和欺骗性物品分类的关键标准。所识别的标准可以体现在自然语言处理系统中以执行非法内容检测。这些标准包括文件的来源和起源、标题、政治观点和几个关键内容特征。本文提出并评估了这些特征的有效性和他们的合法与非法分类的适用性。本文最后讨论使用这些特性作为输入到一个定制的朴素贝叶斯概率分类器,使用这种分类器的结果和未来的工作,其发展。
Fabricated news stories that contain false information but are presented as factually accurate (commonly known as ‘fake news’) have generated substantial interest and media attention following the 2016 U.S. presidential election. While the full details of what transpired during the election are still not known, it appears that multiple groups used social media to spread false information packaged in fabricated news articles that were presented as truthful. Some have argued that this campaign had a material impact on the election. Moreover, the 2016 U.S. presidential election is far from the only campaign where fake news had an apparent role. In this paper, work on a counter-fake-news research effort is presented. In the long term, this project is focused on building an indications and warnings systems for potentially deceptive false content.As part of this project, a dataset of manually classified legitimate and deceptive news articles was curated. The key criteria for classifying legitimate and deceptive articles, identified by the manual classification project, are identified and discussed. The identified criteria can be embodied in a natural language processing system to perform illegitimate content detection. The criteria include the document’s source and origin, title, political perspective, and several key content characteristics. This paper presents and evaluates the efficacy of each of these characteristics and their suitability for legitimate versus illegitimate classification. The paper concludes by discussing the use of these characteristics as input to a customized naïve Bayesian probability classifier, the results of the use of this classifier and future work on its development.