课题基金 / 基金详情

ITR/PE: AVENUE: Adaptable Voice Translation for Minority Languages

ITR/PE: AVENUE: Adaptable Voice Translation for Minority Languages
ITR/PE:AVENUE:针对少数民族语言的自适应语音翻译
批准号:
0121631
负责人:
Jaime Carbonell
金额:
$250.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-09-15 至 2007-08-31

项目摘要

项目成果

Jaime Carbonell的其他基金

相似基金

相关文献

中文摘要
翻译
我们的主要研究目标是开发一个原型的语音翻译通信器,它将为少数民族语言提供跨越语言鸿沟的信息服务,以便允许不同语言的远程用户直接与互联网内容和数据库进行交流,更重要的是与其他使用不同语言的人交流。后者将使信息、教育以及例如卫生服务能够接触到偏远的少数民族语言社区。要实现这一目标,需要在翻译机器学习和跨语言识别方面取得重大进展,以适应更广泛的语言现象。传统的基于迁移规则的机器翻译需要长达一个人-世纪的时间来建立和完善新的语言对。基于统计和示例的机器翻译用大量的双语训练数据取代了人类的编码工作,而这些数据对于大多数少数民族语言来说几乎是无法获得的。如果没有根本的进步,开发时间没有一个数量级的改善,唯一在商业上合理的机器翻译应用程序涉及主要的欧洲语言,日语、中文、韩语、阿拉伯语,也许还有一些相对流行的语言。目前绝大多数人类语言被归入众所周知的机器翻译尘埃堆。我们提出了基于扩展的和新的机器学习方法的新的机器翻译方法。第一种方法是从较少数量级的训练数据中学习的统计机器翻译方法,通过结合指数(最大熵)模型的源-通道联合建模方法,可以更有效地融合先验语言信息(包括词典、词类和已知语言规则类或约束);第二种方法是一种从本地信息者那里获取高质量机器翻译转移规则的新方法,减少了对人类专家的依赖,减少了开发时间。第三种方法基于受控双语语料库和交互式工具从母语信息者那里获取信息,通过一种新的局部约束的种子版本空间方法来推广语义条件的转换规则;第三种方法建立跨多语言的通用语音模型用于语音识别,并在最少的新语言训练数据的情况下使识别器适应新语言。所有这些方法都基于新的和现有的机器学习算法,这些算法将先验知识与有限的新数据相结合,以便在工作中的机器翻译和语音识别和合成系统中快速收敛。主要的社会影响将是对全球信息民主化的重大贡献,这一过程需要弥合当前的语言障碍,特别是对于低密度或经济不利的语言。此外,通过新的语言和声学知识以及现有的教学软件,濒危语言的保存和教学将直接成为可能。如果成功,Avenue(Adaptable Voice-EnabledNatural-Translator for Universal Empower)将成为机器翻译系统的原型,该系统将使全球范围内能够获得多语言信息。
英文摘要
Our primary research goal is to develop a prototype voice-enabledtranslating communicator which will deliver information services acrossthe linguistic divide for minority languages in order allow remotelinguistically-diverse users to communicate directly with Internetcontent and databases, and more importantly to communicate with othersspeaking a different language from their own. The latter will enableinformation, education, and, for example, health services, to reachremote minority-language communities. Achieving this goal requiresmajor advances in machine learning for translation and in cross-languagespeech-recognition adaptability to wider language phenomena.Traditional transfer-rule-based MT requires up to a person-centuryto build and perfect a new language pair. Statistical andExample-Based MT replaces human coding effort by vast amounts ofbilingual training data, which are virtually unobtainable for mostminority languages. Without a radical advance, leading to anover-an-order-of-magnitude improvement in development time, theonly commercially justifiable MT applications involve the majorEuropean languages, Japanese, Chinese, Korean, Arabic and perhaps acouple more relatively-popular languages. The vast majority of humanlanguages are currently relegated to the proverbial MT dust heap.We propose new MT approaches based on extended and new machinelearning methods. The first approach consists of statistical MTmethods that learn from orders of magnitude less training data,and that can more effectively incorporate prior linguistic information(including dictionaries, word classes, and known linguisticrule classes or constraints) by using the joint source-channelmodeling approach combined with exponential (maximum entropy) models.The second approach is a new method for acquiring high-quality MTtransfer rules from native informants which decreases dependence onhuman experts and reduces development time. Semantically-conditionedtransfer rules are generalized via a new locally-constrainedSeeded Version-Space method based on a controlled bilingual corpusand interactive tools to elicit information from native informants.The third method builds general phone models across multiple languagefamilies for speech recognition and adapts the recognizer to newlanguages with minimal new- language training data. All of thesemethods are based on new and existing machine learning algorithmsthat combine prior knowledge with limited amounts of new data inorder to converge quickly on working machine translation and speechrecognition and synthesis systems.The primary societal impact will be a significant contributionto the global democratization of informa- tion, a process thatrequires bridging current linguistic barriers, especially forlow-density or economically- disadvantaged languages. Additionally,preservation and teaching of endangered languages will be directlyenabled by the new linguistic and acoustic knowledge coupled withexisting tutorial software. If successful, Avenue (Adaptable Voice-EnabledNatural-translator for Universal Empowerment) will be the prototypeof an MT system that will empower world-wide access to multilingualinformation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
EAGER: Distributed Learning in Expert Referral Networks
  • 批准号:
    1649225
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.0万
  • 财政年份:
    2016
  • 负责人:
    Jaime Carbonell
  • 依托单位:
EAGER: TEACHER: A Pilot Study on Mining the Web for Customized Curriculum Planning
  • 批准号:
    1350364
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.96万
  • 财政年份:
    2013
  • 负责人:
    Jaime Carbonell
  • 依托单位:
RI: Medium: Interactive Transfer Learning in Dynamic Environments
  • 批准号:
    1065251
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $104.82万
  • 财政年份:
    2011
  • 负责人:
    Jaime Carbonell
  • 依托单位:
LETRAS: A Learning-based Framework for Machine Translation of Low Resource Languages
  • 批准号:
    0534217
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2006
  • 负责人:
    Jaime Carbonell
  • 依托单位:
国内基金
海外基金
纳米粘土增强PE膜性能工艺研发
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    陈梓
  • 依托单位:
基于可解释深度学习的PE2/PE3编辑效率预测研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    刘峰
  • 依托单位:
SET7通过调控糖酵解和氧化还原稳态参与PE发生发展的作用及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    唐金花
  • 依托单位:
结核分枝杆菌PE_PGRS62通过靶向IRF3调 控I型干扰素通路的机制研究