课题基金 / 基金详情

Arabic Dialect Modeling for Speech and Natural Language Processing

Arabic Dialect Modeling for Speech and Natural Language Processing
用于语音和自然语言处理的阿拉伯方言建模
批准号:
0329163
负责人:
Owen Rambow
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-01 至 2006-08-31

项目摘要

项目成果

Owen Rambow的其他基金

相似基金

相关文献

中文摘要
翻译
阿拉伯语实际上是方言的集合,具有重要的语音、形态、词汇和句法差异。然而,在整个阿拉伯世界,标准的书面语言是相同的,现代标准阿拉伯语(MSA),也用于一些官方口头交流(新闻广播,议会辩论)。MSA以古典阿拉伯语为基础,本身不是母语。这种情况对阿拉伯语自动语音识别(ASR)和自然语言处理(NLP)产生了重要的负面影响:由于口语方言不是正式书面的,因此获得足够的语料库来训练当今常用的ASR和NLP工具是昂贵的,例如,用于ASR的语言模型。经验表明,使用MSA文本作为语言模型对于提高方言ASR是无效的。该项目旨在设计一种方法,分层地指定一组密切相关的语言/方言的词法和语法。种子语言(MSA)的语法从现有的树库中自动提取,并根据需要手动增强,而相关方言的语法是手动指定的,直到它与其他方言的语法不同。形态学和词典也采用了类似的方法。然后使用这些正式规范在方言中派生换能器。为了测试这种方法的实用性,使用换能器将MSA语料库转换为(近似)方言文本。这些“创建”的语料库反过来用于训练方言的语言模型,期望在仅使用小型方言语料库或(大型)MSA语料库进行语言建模的基础上提高方言ASR。该项目有可能提高阿拉伯语方言ASR的质量,更广泛地说,增加对密切相关的语言如何正式建模的理解。为阿拉伯语方言开发的自然语言处理工具将提供给研究界。
英文摘要
The Arabic language is actually a collection of dialects with important phonological, morphological, lexical, and syntactic differences. However, throughout the Arab world, the standard written language is the same, Modern Standard Arabic (MSA), that is also used in some official spoken communication (newscasts, parliamentary debates). MSA is based on Classical Arabic and is itself not a native spoken language. This situation has important negative consequences for Arabic automatic speech recognition (ASR) and natural language processing (NLP): since the spoken dialects are not officially written, it is costly to obtain adequate corpora to use for training the kind of ASR and NLP tools commonly in use today, for example, language models for ASR. Experience has shown that using MSA text for language models is ineffective in improving dialect ASR.This project aims at devising a way to hierarchically specify the morphology and syntax of a group of closely related languages/dialects. The syntax of a seed language (MSA) is automatically extracted from an existing treebank and is augmented by hand as needed, while the syntax of related dialects is specified manually to the extent that it differs from that of other dialects. A similar approach is pursued for morphology and the lexicon. These formal specifications are then used to derive transducers among the dialects. To test the utility of this approach, the transducers are used to convert MSA corpora to (an approximation of) dialect text. These "created" corpora in turn are used to train language models for the dialect, with the expectation of improving dialect ASR over the baseline in which only small dialect corpora or (large) MSA corpora are used for language modeling.The project has the potential to improve the quality of ASR for Arabic dialects and, more generally, to increase the understanding of how closely related languages can be modeled formally. The developed NLP tools for Arabic dialects will be made available to the research community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
  • 批准号:
    0751089
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.3万
  • 财政年份:
    2008
  • 负责人:
    Owen Rambow
  • 依托单位:
RI: Email, Social Networks, and Organizations: Investigating How We Use Language to Create and Navigate Social and Organizational Relations
  • 批准号:
    0713548
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $31.57万
  • 财政年份:
    2007
  • 负责人:
    Owen Rambow
  • 依托单位:
CRI: CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian Application
  • 批准号:
    0708183
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.64万
  • 财政年份:
    2007
  • 负责人:
    Owen Rambow
  • 依托单位:
NSF-NATO POSTDOCTORAL FELLOWSHIP
  • 批准号:
    9353729
  • 项目类别:
    Fellowship Award
  • 资助金额:
    $0.0万
  • 财政年份:
    1994
  • 负责人:
    Owen Rambow
  • 依托单位:
海外基金