Examination of classifying hoaxes over SNS using Bayesian Network

Examination of classifying hoaxes over SNS using Bayesian Network
复制标题

使用贝叶斯网络检查 SNS 上的恶作剧分类

DOI:
10.1109/candar.2017.103
复制
发表时间:
2018
期刊:
Procedings of CANDAR 2017: The Fifth International Symposium on Computing and Networking
影响因子:
--
通讯作者:
Michio Sonoda and Jinhui Chao
Michio Sonoda and Jinhui Chao
中科院分区:
--
文献类型:
--
作者:
Ryutaro Ushigome;Takeshi Matsuda;Michio Sonoda and Jinhui Chao

文献摘要

相似文献

预计在 SNS 上发布的短信和数据将广泛用于营销和紧急情况。然而,由于这些内容也有可能是恶作剧,因此有必要建立方法来辨别其真实性。为此,基于对句子中单词频率的分析,提出了词袋模型。然而,由于数据存在偏差,仅靠这种方法似乎不足以对帖子内容进行准确分类。在这项研究中,我们使用贝叶斯网络(称为随机图模型之一)来考虑单词的频率和关系。提出了一种基于日语语义方向词典将内容分类为恶作剧或正常的方法。这种方法的一个潜在问题是字典中的单词数量非常大,可能很难使用所有单词构建贝叶斯网络。在这项研究中,我们将字典中用于分类的单词分为不同的组。这种分组使得在可访问的时间内构建贝叶斯网络成为可能。仿真结果表明,所得到的贝叶斯网络对内容进行分类的准确率较高。
Text message and data posted on SNS are expected wide usage for marketing and in occasions of emergency. However, since there is also possibility that these contents are hoaxes, it is necessary to establish methods in order to discriminate their authenticity. A Bag of Words model has been proposed for this purpose based on analysis of frequency of words in sentences. However, it seemed that this method alone may not be enough to classify post contents accurately due to bias in the data. In this research, we use the Bayesian network, known as one of stochastic graphical models, in order to take into account of both the frequency and the relation of words. A method is proposed to classify a content as either a hoax or normal based on Japanese semantic orientation dictionary. A potential problem of this approach is that the number of words in the dictionary is very large, it may be difficult to construct a Bayesian network using all of words. In this research, we separate words in the dictionary used for classification into different groups. This grouping made it possible to construct a Bayesian network in accessible time. Simulations shown that the obtained Bayesian network classifies contents with high accuracy.