Learning from Relatives: Unified Dialectal Arabic Segmentation

Learning from Relatives: Unified Dialectal Arabic Segmentation
复制标题

DOI:
10.18653/v1/k17-1043
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Younes Samih;Mohamed I. Eldesouki;Mohammed Attia;Kareem Darwish;Ahmed Abdelali;Hamdy Mubarak;Laura Kallmeyer
Younes Samih;Mohamed I. Eldesouki;Mohammed Attia;Kareem Darwish;Ahmed Abdelali;Hamdy Mubarak;Laura Kallmeyer
中科院分区:
其他
文献类型:
--
作者:
Younes Samih;Mohamed I. Eldesouki;Mohammed Attia;Kareem Darwish;Ahmed Abdelali;Hamdy Mubarak;Laura Kallmeyer

文献摘要

被引文献

相似文献

阿拉伯语方言不仅共享一个共同的koiné,而且还有共享的泛方言语言现象,允许方言的计算模型相互学习。在本文中,我们建立了一个统一的分割模型,其中不同方言的训练数据相结合,并训练一个单一的模型。该模型产生更高的准确性比方言特定的模型,消除了方言识别之前分割的需要。我们还通过测试在一种方言上训练的分割模型在其他方言上的表现来衡量四种主要阿拉伯语方言之间的相关程度。我们发现,语言的相关性是偶然的地理接近。在我们的实验中,我们使用基于SVM的排名和bi-LSTM-CRF序列标记。
Arabic dialects do not just share a common koiné, but there are shared pan-dialectal linguistic phenomena that allow computational models for dialects to learn from each other. In this paper we build a unified segmentation model where the training data for different dialects are combined and a single model is trained. The model yields higher accuracies than dialect-specific models, eliminating the need for dialect identification before segmentation. We also measure the degree of relatedness between four major Arabic dialects by testing how a segmentation model trained on one dialect performs on the other dialects. We found that linguistic relatedness is contingent with geographical proximity. In our experiments we use SVM-based ranking and bi-LSTM-CRF sequence labeling.