Improving Sinhala Hate Speech Detection Using Deep Learning
Improving Sinhala Hate Speech Detection Using Deep Learning
复制标题
使用深度学习改进僧伽罗仇恨言论检测
DOI:
10.1109/icter58063.2022.10024103
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
R. Weerasinghe
中科院分区:
文献类型:
--
作者:
K. Gamage;V. Welgama;R. Weerasinghe
Automatic Hate Speech Detection is a fine-grained sentiment analysis task that has been the focus of many researchers around the world. This has been a difficult task due to challenges such as the usage of native languages and distinct vocabularies, as well as the distortion of words. However, based on the findings of previous studies on Sinhala hate speech identification, this has proven to be more difficult for low-resource languages like Sinhala. The effectiveness of pretrained embedding for Sinhala hate speech detection has not been investigated. We investigated several embeddings as well as frequency-based features, including bag of words, n-grams, and TF-IDF to address this shortcoming. We present results from several machine learning experiments, including deep learning experiments and transfer learning experiments on state-of-the-art cross-lingual transformers. With an f1-score of 0.764 and a recall value of 0.788 in our study, the XLMR model outperformed other baseline algorithms and deep learning models.