AraSenCorpus: A Semi-Supervised Approach for Sentiment Annotation of a Large Arabic Text Corpus

AraSenCorpus: A Semi-Supervised Approach for Sentiment Annotation of a Large Arabic Text Corpus
复制标题

DOI:
10.3390/app11052434
复制
发表时间:
2021-03-01
影响因子:
2.7
通讯作者:
Rehmat, Asim
Rehmat, Asim
中科院分区:
综合性期刊4区
文献类型:
--
作者:
Al-Laith, Ali;Shahbaz, Muhammad;Rehmat, Asim

文献摘要

被引文献

相似文献

当情感分析领域的研究倾向于研究语言中的高级主题时,如英语,其他语言,如阿拉伯语仍然面临基本问题和挑战,最明显的是大型语料库的可用性。此外,当语料库太大时,人工标注是耗时且困难的。本文提出了一种半监督自学习技术,扩展阿拉伯语情感标注语料库与未标记的数据,命名为AraSenCorpus。我们使用神经网络在包含15,000条推文的手动标记数据集上训练一组模型。我们使用这些模型来扩展语料库的大型阿拉伯语情感语料库称为“AraSenCorpus”。AraSenCorpus包含450万条推文,涵盖现代标准阿拉伯语和一些阿拉伯方言。长短期记忆(LSTM)深度学习分类器用于训练和测试最终语料库。我们在两个外部基准数据集上评估了我们提出的框架,以确保阿拉伯语情感分类的改进。实验结果表明,我们的语料库优于现有的国家的最先进的系统。
At a time when research in the field of sentiment analysis tends to study advanced topics in languages, such as English, other languages such as Arabic still suffer from basic problems and challenges, most notably the availability of large corpora. Furthermore, manual annotation is time-consuming and difficult when the corpus is too large. This paper presents a semi-supervised self-learning technique, to extend an Arabic sentiment annotated corpus with unlabeled data, named AraSenCorpus. We use a neural network to train a set of models on a manually labeled dataset containing 15,000 tweets. We used these models to extend the corpus to a large Arabic sentiment corpus called "AraSenCorpus". AraSenCorpus contains 4.5 million tweets and covers both modern standard Arabic and some of the Arabic dialects. The long-short term memory (LSTM) deep learning classifier is used to train and test the final corpus. We evaluate our proposed framework on two external benchmark datasets to ensure the improvement of the Arabic sentiment classification. The experimental results show that our corpus outperforms the existing state-of-the-art systems.