Improving Sinhala Hate Speech Detection Using Deep Learning

Improving Sinhala Hate Speech Detection Using Deep Learning
复制标题

使用深度学习改进僧伽罗仇恨言论检测

DOI:
10.1109/icter58063.2022.10024103
复制
发表时间:
2022
期刊:
2022 22nd International Conference on Advances in ICT for Emerging Regions (ICTer)
影响因子:
--
通讯作者:
R. Weerasinghe
R. Weerasinghe
中科院分区:
--
文献类型:
--
作者:
K. Gamage;V. Welgama;R. Weerasinghe

文献摘要

被引文献

相似文献

自动仇恨语音检测是一个细粒度的情感分析任务,一直是世界各地许多研究人员的焦点。这是一项艰巨的任务,因为存在着诸如使用母语和不同词汇以及单词失真等挑战。然而,根据以前关于僧伽罗语仇恨言论识别的研究结果,这对于像僧伽罗语这样的低资源语言来说更加困难。尚未研究预训练嵌入对僧伽罗仇恨言论检测的有效性。我们研究了几种嵌入以及基于频率的特征,包括词袋,n-gram和TF-IDF来解决这个缺点。我们展示了几个机器学习实验的结果,包括深度学习实验和最先进的跨语言转换器上的迁移学习实验。在我们的研究中,XLMR模型的f1得分为0.764,召回值为0.788,优于其他基线算法和深度学习模型。
Automatic Hate Speech Detection is a fine-grained sentiment analysis task that has been the focus of many researchers around the world. This has been a difficult task due to challenges such as the usage of native languages and distinct vocabularies, as well as the distortion of words. However, based on the findings of previous studies on Sinhala hate speech identification, this has proven to be more difficult for low-resource languages like Sinhala. The effectiveness of pretrained embedding for Sinhala hate speech detection has not been investigated. We investigated several embeddings as well as frequency-based features, including bag of words, n-grams, and TF-IDF to address this shortcoming. We present results from several machine learning experiments, including deep learning experiments and transfer learning experiments on state-of-the-art cross-lingual transformers. With an f1-score of 0.764 and a recall value of 0.788 in our study, the XLMR model outperformed other baseline algorithms and deep learning models.