Machine Learning-Based Approach for Arabic Dialect Identification

Machine Learning-Based Approach for Arabic Dialect Identification
复制标题

基于机器学习的阿拉伯语方言识别方法

DOI:
--
复制
发表时间:
2021
期刊:
Workshop on Arabic Natural Language Processing
影响因子:
--
通讯作者:
A. El
A. El
中科院分区:
--
文献类型:
--
作者:
Hamada Nayel;Ahmed Hassan;Mahmoud Sobhi;A. El

文献摘要

被引文献

相似文献

本文描述了我们提交给第二次细致阿拉伯方言识别共享任务 (NADI 2021) 的系统。方言识别是自动检测给定文本或语音片段的源类型的任务。共有四个子任务,其中两个子任务为国家级识别,另外两个子任务为省级识别。该任务的数据覆盖了所有21个阿拉伯国家的100个省份,并且来自Twitter域。所提出的系统依赖于五种机器学习方法,即补朴素贝叶斯、支持向量机、决策树、逻辑回归和随机森林分类器。朴素贝叶斯分类器的 F1 宏观平均得分在开发和测试数据方面优于所有其他分类器。
This paper describes our systems submitted to the Second Nuanced Arabic Dialect Identification Shared Task (NADI 2021). Dialect identification is the task of automatically detecting the source variety of a given text or speech segment. There are four subtasks, two subtasks for country-level identification and the other two subtasks for province-level identification. The data in this task covers a total of 100 provinces from all 21 Arab countries and come from the Twitter domain. The proposed systems depend on five machine-learning approaches namely Complement Naïve Bayes, Support Vector Machine, Decision Tree, Logistic Regression and Random Forest Classifiers. F1 macro-averaged score of Naïve Bayes classifier outperformed all other classifiers for development and test data.