A novel approach to stance detection in social media tweets by fusing ranked lists and sentiments

A novel approach to stance detection in social media tweets by fusing ranked lists and sentiments
复制标题

DOI:
10.1016/j.inffus.2020.10.003
复制
发表时间:
2021-03-01
期刊:
影响因子:
18.6
通讯作者:
Hussain, Amir
Hussain, Amir
中科院分区:
计算机科学1区
文献类型:
--
作者:
Al-Ghadir, Abdulrahman I.;Azmi, Aqil M.;Hussain, Amir

文献摘要

被引文献

相似文献

立场检测是数据挖掘中一个相对较新的概念,旨在为社交媒体帖子分配一个立场标签(支持,反对或无),以达到特定的预定目标。这些目标可能不会在帖子中提及,也可能不会成为帖子中的意见目标。在本文中,我们提出了一种新颖的增强方法来识别给定推文作者的立场。这包括用于姿态检测的三个阶段的过程:(a)推文预处理;在这里,我们清理和规范化推文(例如,去除停用词)以生成词和词干列表,(B)特征生成;在该步骤中,我们创建并融合两个字典以生成特征向量,以及最后(c)分类;基于目标列表对特征的所有实例进行分类。我们创新的特征选择提出了两个排序列表(前k)的词频逆文档频率(tf-idf)分数和情感信息的融合。我们使用六种不同的分类器来评估我们的方法:K最近邻(K-NN),基于可信度的K-NN,加权K-NN,基于类的K-NN,基于样本的K-NN和支持向量机。此外,我们调查使用主成分分析,并研究其对性能的影响。在基准数据集(SemEval-2016任务6)上评估模型,并使用t检验确定结果的显著性。我们使用加权K-NN分类器实现了76.45%的宏观F分数(所有主题的平均值)的最佳性能。这超过了同一数据集上目前最先进的74.44%的分数。
Stance detection is a relatively new concept in data mining that aims to assign a stance label (favor, against, or none) to a social media post towards a specific pre-determined target. These targets may not be referred to in the post, and may not be the target of opinion in the post. In this paper, we propose a novel enhanced method for identifying the writer's stance of a given tweet. This comprises a three-phase process for stance detection: (a) tweets preprocessing; here we clean and normalize tweets (e.g., remove stop-words) to generate words and stems lists, (b) features generation; in this step, we create and fuse two dictionaries for generating features vector, and lastly (c) classification; all the instances of the features are classified based on the list of targets. Our innovative feature selection proposes fusion of two ranked lists (top -k) of term frequency-inverse document frequency (tf-idf) scores and the sentiment information. We evaluate our method using six different classifiers: K nearest neighbor (K-NN), discernibility-based K-NN, weighted K-NN, class-based K-NN, exemplar-based K-NN, and Support Vector Machines. Furthermore, we investigate the use of Principal Component Analysis and study its effect on performance. The model is evaluated on the benchmark dataset (SemEval-2016 task 6), and the results significance is determined using t-test. We achieve our best performance of macro F-score (averaged across all topics) of 76.45% using the weighted K-NN classifier. This tops the current state-of-the-art score of 74.44% on the same dataset.