课题基金 / 基金详情

ITR: Applying Translation Technology to Language Modeling

ITR: Applying Translation Technology to Language Modeling
ITR:将翻译技术应用于语言建模
批准号:
0326276
负责人:
Mari Ostendorf
金额:
$300.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2003
资助国家:
美国
项目状态:
已结题
起止时间:
2003-09-15 至 2008-08-31

项目摘要

项目成果

Mari Ostendorf的其他基金

相似基金

相关文献

中文摘要
翻译
几乎所有产生文本的系统,从语音识别到自然语言生成,都使用语言模型作为核心组件,以便根据它们的结构良好性和对给定上下文的适当性来对单词串进行排名。这些模型很难开发,既是因为集成多个知识来源所面临的算法挑战,也是因为缺乏强大的语言处理工具。该项目的目标是通过新技术开发模型,以便利用平行多语种语料库中可用的信息,即同一来源的多种语文的翻译。这样的语料库隐含地编码了一个隐藏的、共同的核心,可以使用最先进的估计技术来发现。该项目涉及:i)在多个抽象层次上自动学习语言内部和跨语言的结构:语义、词法、音系学和释义,以及ii)将结果整合到新的语言模型框架中,以解决特定领域和特定语言的训练数据有限的问题。假设是,通过在一种语言内跨语言和流派共享数据和结构,所产生的模型将更加丰富和可靠。这样的想法直到最近才是不可能想象的;多语言语料库的可用性和计算能力的增加使它们现在成为可能。该项目结合了机器翻译和语音识别语言建模技术,预计这两种技术的结合将产生更强大和更通用的模型。这项研究将促进学习程度较低的语言工具的快速开发,并将立即影响主流语言的应用,从信息管理到国际合作再到双语教育。研究结果还将对语言处理以外的统计建模问题产生影响。
英文摘要
Virtually all systems that produce text, from speech recognition to natural language generation, use a language model as a core component in order to rank word strings by their well-formedness and appropriateness for a given context. These models are difficult to develop both because of algorithmic challenges specific to integration of multiple knowledge sources and the lack of robust language processing tools. The goal of this project is to develop models via new techniques for exploiting the information available in parallel multilingual corpora, i.e., translations of the same source in multiple languages. Such corpora implicitly encode a hidden, common core that can be uncovered using state-of-the-art estimation techniques. The project involves: i) automatic learning of structure within and across languages at multiple levels of abstraction: semantics, morphology, phonology, and paraphrasing, and ii) integration of the results into novel language model frameworks to address the problem of limited domain- and language-specific training data. The hypothesis is that, by sharing data and structure across languages and genres within a language, the resulting models will be richer and more robust. Such ideas were impossible to envision until recently; availability of multilingual corpora and increases in computing power make them now feasible.This project marries machine translation and speech recognition language modeling techniques, anticipating that the combination will lead to more powerful and general models. The research will facilitate rapid development of tools for less well studied languages and will immediately impact applications in mainstream languages ranging from information management to international collaboration to bilingual education. The results will also have implications for statistical modeling problems beyond language processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Improving Speech Technology for Better Learning Outcomes: The Case of AAE Child Speakers
  • 批准号:
    2202049
  • 项目类别:
    Standard Grant
  • 资助金额:
    $26.01万
  • 财政年份:
    2022
  • 负责人:
    Mari Ostendorf
  • 依托单位:
RI: Small: Modeling Idiosyncrasies of Speech for Automatic Spoken Language Processing
  • 批准号:
    1617176
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2016
  • 负责人:
    Mari Ostendorf
  • 依托单位:
RI: Small: Simplifying Text for Individual Reading Needs
  • 批准号:
    0916951
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.0万
  • 财政年份:
    2009
  • 负责人:
    Mari Ostendorf
  • 依托单位:
U.S.-Germany Dissertation Enhancement: Predicting Hidden Structure and Punctuation in Speech for Machine Translation
  • 批准号:
    0552492
  • 项目类别:
    Standard Grant
  • 资助金额:
    $1.5万
  • 财政年份:
    2006
  • 负责人:
    Mari Ostendorf
  • 依托单位:
海外基金