Sexism detection: The first corpus in Algerian dialect with a code-switching in Arabic/ French and English

Sexism detection: The first corpus in Algerian dialect with a code-switching in Arabic/ French and English
复制标题

性别歧视检测:第一个具有阿拉伯语/法语和英语语码转换功能的阿尔及利亚方言语料库

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Akram Abdelhaq Moumna
Akram Abdelhaq Moumna
中科院分区:
--
文献类型:
--
作者:
I. Guellil;A. Adeel;F. Azouaou;Mohamed Boubred;Yousra Houichi;Akram Abdelhaq Moumna

文献摘要

被引文献

相似文献

本文提出了一种针对阿拉伯社区女性在社交媒体(如Youtube)上的仇恨言论检测方法。在文献中,类似的作品已经提出了其他语言,如英语。然而,据我们所知,以阿拉伯文进行的工作并不多。使用三种不同的注释器开发了一个新的仇恨言论语料库(Arabic\_fr\_en)。对于语料库验证,使用了三种不同的机器学习算法,包括深度卷积神经网络(CNN)、长短期记忆(LSTM)网络和双向LSTM (Bi-LSTM)网络。仿真结果表明,与LSTM和Bi-LSTM相比,CNN模型在不平衡语料库上的f1得分高达86%。
In this paper, an approach for hate speech detection against women in Arabic community on social media (e.g. Youtube) is proposed. In the literature, similar works have been presented for other languages such as English. However, to the best of our knowledge, not much work has been conducted in the Arabic language. A new hate speech corpus (Arabic\_fr\_en) is developed using three different annotators. For corpus validation, three different machine learning algorithms are used, including deep Convolutional Neural Network (CNN), long short-term memory (LSTM) network and Bi-directional LSTM (Bi-LSTM) network. Simulation results demonstrate the best performance of the CNN model, which achieved F1-score up to 86\% for the unbalanced corpus as compared to LSTM and Bi-LSTM.