Machine Learning-Based Approach for Arabic Dialect Identification
Machine Learning-Based Approach for Arabic Dialect Identification
复制标题
基于机器学习的阿拉伯语方言识别方法
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
A. El
中科院分区:
文献类型:
--
作者:
Hamada Nayel;Ahmed Hassan;Mahmoud Sobhi;A. El
This paper describes our systems submitted to the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). Dialect identification is the task of automatically detecting the source variety of a given text or speech segment. There are four subtasks, two subtasks for country-level identification and the other two subtasks for province-level identification. The data in this task covers a total of 100 provinces from all 21 Arab countries and come from the Twitter domain. The proposed systems depend on five machine-learning approaches namely Complement Naïve Bayes, Support Vector Machine, Decision Tree, Logistic Regression and Random Forest Classifiers. F1 macro-averaged score of Naïve Bayes classifier outperformed all other classifiers for development and test data.