Collaborative Research: Structure Alignment-based Machine Translation
Collaborative Research: Structure Alignment-based Machine Translation
批准号:
0534325
负责人:
Michiko Kosaka
金额:
$8.2万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2006
资助国家:
美国
项目状态:
已结题
起止时间:
2006-07-01 至 2010-06-30
中文摘要
纽约大学、蒙茅斯大学和科罗拉多大学的研究人员正在构建日语/英语和汉语/英语机器翻译系统,该系统可以自动从对平行文本的“深度”语言分析中获取规则。这项工作是自动化基于示例的机器翻译(MT)项目的自然高潮,该项目在过去二十年中变得越来越复杂。自然语言处理(NLP)技术的以下最新进展使这种查询变得可行:(1)注释数据,包括双语树库和在此数据上训练的处理器(解析器,PropBankers等);(2)解析器输出的语义后处理器;(3)自动对齐二进制文本的程序;(4)双语树到树翻译模型。自然语言在跨语言边界和同一种语言的等价表达式的对应词的顺序上差别很大。本研究探讨了使用一种自动从语法树中派生出来的语义表示(GLARF)来最小化单一语言中的差异的方法。这种语义表示提供了:(1)减少表示相同底层消息的方法的数量,以及(2)将长距离依赖关系(例如相对子句)作为本地现象处理的方法。因此,没有必要借助任意长的句子片段或大树进行训练。此外,由于所需的数据较少,它最大限度地减少了稀疏数据问题。在这个翻译模型的训练中,由于(1),源树和目标树之间的映射规则的数量减少了。因此,翻译模型是一个树形换能器,具有对源表示和目标表示进行“深度”语言分析的树形。为了为这种部分映射提供高效的计算机算法,本研究需要重点关注(a)训练算法和(b)映射规则的约束,以降低计算复杂度。这项研究有望产生几个优势:使用“深度”语言分析的换能器的核心架构应该产生更准确的结果。GLARF架构允许控制不同粒度的自动获得的语言分析。更广泛的影响:对机器翻译的需求涵盖了从地方政府(例如警察部队)到国家政府(例如政府)。中央情报局)和私营部门。鉴于互联网在英语世界之外的发展,更好的机器翻译对更广泛的社区至关重要。这项工作直接影响到英语使用者理解中文和日文网站的能力,而中文和日文是互联网上使用最广泛的两种语言。该技术可推广到其他语言对,并最终产生更广泛的影响。
英文摘要
Researchers at New York University, Monmouth University and theUniversity of Colorado are constructing Japanese/English andChinese/English machine translation systems which automatically acquirerules from ``deep'' linguistic analyses of parallel text. This workis a natural culmination of automated example-based MachineTranslation (MT) projects that have become increasingly sophisticatedover the last two decades. The following recent advances in NaturalLanguage Processing (NLP) technologies make this inquiry feasible: (1)annotated data including bilingual treebanks and processors trained onthis data (parsers, PropBankers, etc.); (2) semantic post-processorsof parser output; (3) programs that automatically align bitexts; and(4) bilingual tree to tree translation models.Natural languages vary widely in the ordering of corresponding wordsfor equivalent expressions across linguistic boundaries and within asingle language. This research investigates ways to minimize thevariations within a single language using a type of semanticrepresentation (GLARF) that is derived automatically from syntactictrees. Such semantic representation provides for: (1) a reduction in thenumber of ways of representing the same underlying message, and (2)a way to handle long distance dependencies (e.g. relativeclauses) as local phenomena. Therefore, there is no need to resort toarbitrarily long sentence fragments or large trees fortraining. Furthermore, since less data is needed, itminimizes the sparse data problem.In the training of this translation model, because of (1), the numberof mapping rules between the source tree and the target tree isreduced. The translation model, then, is a tree transducer, with``deep'' linguistically analyzed trees for both source and targetrepresentations. In order to provide efficient computer algorithmsfor such partial mappings, this research needs to focus on(a) the training algorithm and the (b) the constraints over themapping rules in order to reduce the computational complexity.This research is expected to yield several advantages: The corearchitecture of this transducer using ``deep'' linguistic analysesshould yield more accurate results. The GLARF architecture allowscontrol over different granularity of automatically-obtainedlinguistic analyses.Broader Impact: The demand for machine translation spans from thelocal government (e.g. police forces) to national government(e.g. CIA) and the private sector. Given the growth of the Internetoutside the English speaking world, better machine translation is ofcritical importance for the broader community. This work directlyaffects the ability of English speakers to understand websites writtenin Chinese and Japanese, two of the most widely used languages on theInternet. The technique is generalizable to other language pairs andcan ultimately have even wider impact.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
"Collaborative Research on Knowledge Aquisition for Japanese-English Machine Translation"
-
批准号:9302903
-
项目类别:Continuing Grant
-
资助金额:$22.4万
-
财政年份:1993
-
负责人:Michiko Kosaka
-
依托单位:
A Sublanguage Approach to Japanese-English Machine Translation
-
批准号:8902269
-
项目类别:Continuing Grant
-
资助金额:$10.61万
-
财政年份:1989
-
负责人:Michiko Kosaka
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: