Parsing Arabic Dialects

Parsing Arabic Dialects
复制标题

解析阿拉伯语方言

DOI:
10.7916/d85t3tzg
复制
发表时间:
2006
期刊:
影响因子:
2.1
通讯作者:
S. Shareef
S. Shareef
中科院分区:
人文科学1区
文献类型:
--
作者:
David Chiang;Mona T. Diab;Nizar Habash;Owen Rambow;S. Shareef

文献摘要

被引文献

相似文献

阿拉伯语是口语方言的集合,具有重要的语音、形态、词汇和句法差异,以及标准书面语言,现代标准阿拉伯语 (MSA)。由于口语方言不是正式书写的,因此获取足够的语料库来训练方言 NLP 工具(例如解析器)的成本非常高。在本文中,我们解决了解析转录的黎凡特阿拉伯语口语 (LA) 的问题。我们不假设存在任何带注释的 LA 语料库(开发和测试除外),也不假设存在平行语料库 LAMSA。相反,我们使用有关 LA 和 MSA 之间关系的显性知识。
The Arabic language is a collection of spoken dialects with important phonological, morphological, lexical, and syntactic differences, along with a standard written language, Modern Standard Arabic (MSA). Since the spoken dialects are not officially written, it is very costly to obtain adequate corpora to use for training dialect NLP tools such as parsers. In this paper, we address the problem of parsing transcribed spoken Levantine Arabic (LA).We do not assume the existence of any annotated LA corpus (except for development and testing), nor of a parallel corpus LAMSA. Instead, we use explicit knowledge about the relation between LA and MSA.