A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic

A Multi-Dialect, Multi-Genre Corpus of Informal Written Arabic
复制标题

多方言、多流派的非正式书面阿拉伯语语料库

DOI:
--
复制
发表时间:
2014
期刊:
International Conference on Language Resources and Evaluation
影响因子:
--
通讯作者:
Chris Callison
Chris Callison
中科院分区:
--
文献类型:
--
作者:
Ryan Cotterell;Chris Callison

文献摘要

被引文献

相似文献

本文介绍了一个多方言、多体裁、人工标注的方言阿拉伯语语料库。我们收集了五种阿拉伯方言的话语:黎凡特方言、海湾方言、埃及方言、伊拉克方言和马格里比方言。我们在报纸网站上搜索用户评论,在Twitter上搜索两种不同类型的方言内容。据作者所知,在内容来源和方言数量上,这部作品都是最多样化的阿拉伯语方言语料库。语料库中的每一句话都是在亚马逊机械突厥语上进行人工标注的;这与Al-Sabbagh和Girju(2012)形成了鲜明对比,后者只对语料库的其余部分进行了人工标注,以便训练分类器自动标注语料库的其余部分。除了单个工作者的表现之外,我们还提供了用于注释的方法的讨论。我们将阿拉伯语方言识别任务扩展到伊拉克方言和马格里比方言,改进了Zaidan和Callison-Burch(2011a)对Levine、Bay和埃及方言的识别结果。
This paper presents a multi-dialect, multi-genre, human annotated corpus of dialectal Arabic. We collected utterances in five Arabic dialects: Levantine, Gulf, Egyptian, Iraqi and Maghrebi. We scraped newspaper websites for user commentary and Twitter for two distinct types of dialectal content. To the best of the authors knowledge, this work is the most diverse corpus of dialectal Arabic in both the source of the content and the number of dialects. Every utterance in the corpus was human annotated on Amazons Mechanical Turk; this stands in contrast to Al-Sabbagh and Girju (2012) where only a small subset was human annotated in order to train a classifier to automatically annotate the remainder of the corpus. We provide a discussion of the methodology used for the annotation in addition to the performance of the individual workers. We extend the Arabic dialect identification task to the Iraqi and Maghrebi dialects and improve the results of Zaidan and Callison-Burch (2011a) on Levantine, Gulf and Egyptian.