课题基金 / 基金详情

Arabic Dialect Modeling for Speech and Natural Language Processing

Arabic Dialect Modeling for Speech and Natural Language Processing
用于语音和自然语言处理的阿拉伯方言建模
批准号:
0329163
负责人:
Owen Rambow
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-01 至 2006-08-31

项目摘要

项目成果

Owen Rambow的其他基金

相似基金

相关文献

中文摘要
翻译
阿拉伯语实际上是具有重要的语音、形态、词汇和句法差异的方言的集合。然而,在整个阿拉伯世界,标准的书面语言是相同的,现代标准阿拉伯语(MSA),也用于一些官方口头交流(新闻广播,议会辩论)。MSA以古典阿拉伯语为基础,本身不是一种母语。这种情况对阿拉伯语自动语音识别(ASR)和自然语言处理(NLP)有重要的负面影响:由于口语方言不是正式书写的,因此要获得足够的语料库来培训今天常用的ASR和NLP工具,例如ASR的语言模型,成本很高。经验表明,将MSA文本用于语言模型在改进方言ASR方面是无效的。该项目旨在设计一种方法来分层地指定一组密切相关的语言/方言的词法和句法。种子语言(MSA)的语法是从现有树库中自动提取的,并根据需要手动增加,而相关方言的语法是手动指定的,其程度与其他方言的不同。词法和词汇也采用了类似的方法。然后,这些正式规范被用来在各方言之间派生传感器。为了测试这种方法的有效性,使用传感器将MSA语料库转换为方言文本(近似)。这些“创建的”语料库反过来被用来训练方言的语言模型,以期在只使用小的方言语料库或(大型)MSA语料库进行语言建模的基础上改进方言的ASR。该项目有可能提高阿拉伯方言的ASR的质量,更广泛地说,增加对密切相关的语言可以如何正式建模的理解。将向研究界提供为阿拉伯方言开发的自然语言处理工具。
英文摘要
The Arabic language is actually a collection of dialects with important phonological, morphological, lexical, and syntactic differences. However, throughout the Arab world, the standard written language is the same, Modern Standard Arabic (MSA), that is also used in some official spoken communication (newscasts, parliamentary debates). MSA is based on Classical Arabic and is itself not a native spoken language. This situation has important negative consequences for Arabic automatic speech recognition (ASR) and natural language processing (NLP): since the spoken dialects are not officially written, it is costly to obtain adequate corpora to use for training the kind of ASR and NLP tools commonly in use today, for example, language models for ASR. Experience has shown that using MSA text for language models is ineffective in improving dialect ASR.This project aims at devising a way to hierarchically specify the morphology and syntax of a group of closely related languages/dialects. The syntax of a seed language (MSA) is automatically extracted from an existing treebank and is augmented by hand as needed, while the syntax of related dialects is specified manually to the extent that it differs from that of other dialects. A similar approach is pursued for morphology and the lexicon. These formal specifications are then used to derive transducers among the dialects. To test the utility of this approach, the transducers are used to convert MSA corpora to (an approximation of) dialect text. These "created" corpora in turn are used to train language models for the dialect, with the expectation of improving dialect ASR over the baseline in which only small dialect corpora or (large) MSA corpora are used for language modeling.The project has the potential to improve the quality of ASR for Arabic dialects and, more generally, to increase the understanding of how closely related languages can be modeled formally. The developed NLP tools for Arabic dialects will be made available to the research community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: CRI: CRD: A Multi-Representational and Multi-Layered Treebank for Hindi/Urdu
  • 批准号:
    0751089
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $12.3万
  • 财政年份:
    2008
  • 负责人:
    Owen Rambow
  • 依托单位:
RI: Email, Social Networks, and Organizations: Investigating How We Use Language to Create and Navigate Social and Organizational Relations
  • 批准号:
    0713548
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $31.57万
  • 财政年份:
    2007
  • 负责人:
    Owen Rambow
  • 依托单位:
CRI: CRD Collaborative Research: General Techniques for Creating Treebanks with Multiple Representations: A Large-Scale Russian Application
  • 批准号:
    0708183
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.64万
  • 财政年份:
    2007
  • 负责人:
    Owen Rambow
  • 依托单位:
NSF-NATO POSTDOCTORAL FELLOWSHIP
  • 批准号:
    9353729
  • 项目类别:
    Fellowship Award
  • 资助金额:
    $0.0万
  • 财政年份:
    1994
  • 负责人:
    Owen Rambow
  • 依托单位:
海外基金