课题基金 / 基金详情

Models of morphosyntax for statistical machine translation

Models of morphosyntax for statistical machine translation
统计机器翻译的形态句法模型
批准号:
123083856
负责人:
Professor Dr. Hinrich Schütze
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2009
资助国家:
德国
项目状态:
已结题
起止时间:
2008-12-31 至 2017-12-31

项目摘要

项目成果

Professor Dr. Hinrich Schütze的其他基金

相似基金

相关文献

中文摘要
翻译
在过去的几年里,机器翻译(MT)的统计方法已经证明了自己是有效的。然而,当翻译成一种形态丰富的语言时,这是不正确的,特别是当两种语言之间也存在显著的句法差异时。在这种情况下,统计机器翻译(SMT)的质量很差,因为形态学、句法和翻译模型之间的独立假设不能反映语言现实。在项目的第一阶段,我们在翻译成德语方面取得了重大进展,德语是一种形态丰富的语言。重点研究了统计机器翻译中的语言表示和语言资源问题。我们开展了关于德语构词法(包括复合词和合成词)、德语屈折形态和语法问题的原创性研究,这些问题涉及英语到德语的翻译和德语到英语的翻译。我们在顶级国际会议上发表了7篇会议出版物,并发表了2篇研讨会论文,并指导了与该项目相关的本科和硕士学生工作。在提议的第二阶段,我们将从关注语言表示转向研究先进的机器学习方法,以解决困难的机器翻译语言对英语/德语中固有的语言问题。在前一阶段的工作中,我们主要关注翻译中的一般语言问题。在分析语言增强系统的输出时,我们学到的最重要的教训之一是,训练数据和测试数据之间的域不匹配问题是一个非常重要的问题。训练数据主要来自欧洲议会会议记录,但测试数据来自新闻领域或许多其他领域(包括医学领域,我们将在项目的这个阶段学习)。在拟建的morphsyntax项目的第二阶段,我们将遵循四条主要工作路线。我们将扩展我们在德语构词和屈折形态学方面的成功工作,通过确定如何进行半监督的形态学资源获取来减少我们对手工制作的形态学资源的依赖,特别关注领域适应的重要问题。在分层解码(语法形式被用作翻译的表示)中,我们将研究高级机器学习方法的集成,以选择语法重排序。我们还将研究在分层解码中添加语义信息(除了语法信息外)的问题。最后,我们将研究将强大的分类方法(其他工作包所需)直接集成到解码器中的不同方法,而不是使用外部预处理或后处理。
英文摘要
Statistical approaches to machine translation (MT) have shownthemselves to be effective in the last few years. However, whentranslating into a morphologically rich language this is not true,particularly when there is also significant syntactic divergencebetween the two languages. The quality of statistical machinetranslation (SMT) is poor in this case because of independenceassumptions made between the models of morphology, syntax andtranslation that do not reflect linguistic reality.In the first phase of the project we made significant strides intranslating into German, a morphologically rich language. We focusedon issues of linguistic representation and linguistic resources withinstatistical machine translation. We carried out original research indealing with German word formation (addressing both compounds andportmanteaus), German inflectional morphology and syntactic issues indealing with both English to German translation and German to Englishtranslation. We published seven conference publications at top rankedinternational conferences as well as two workshop contributions, andalso supervised Bachelors-level and Masters-level student workrelevant to the project.In the proposed phase 2, we will move on from afocus on linguistic representation to working on advanced machinelearning approaches for solving the linguistic problems inherent inthe difficult machine translation language pair English/German. In theprevious phase of the work, we focused on general linguistic problemsin translation. One of the most important lessons we learned inanalyzing the output of our linguistically enhanced systems is thatthe issue of the mismatch of domains between training data and testdata is a critically important issue. The training data is mostlytaken from the European parliament proceedings, but the testing datais from the news domain or many other domains (including the medicaldomain, which we will study in this phase of the project).In the proposed phase 2 of the Morphosyntax project, we will followfour main lines of work. We will extend our successful work on Germanword formation and inflectional morphology by reducing our dependenceon hand-crafted morphological resources by determining how to performsemi-supervised acquisition of morphological resources, payingparticular attention to the important issue of domainadaptation. Within hierarchical decoding (wheresyntactic formalisms are used as the representation for translation),we will study the integration of advanced machine learning methods forchoosing syntactic reorderings. We will also study the issue of addingsemantic information (in addition to syntactic information) intohierarchical decoding. Finally, we will study different ways tointegrate powerful classification approaches (required for the otherwork packages) directly into the decoder rather than using externalpre-processing or post-processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ReMLAV: Relational Machine Learning for Argument Validation
  • 批准号:
    376183703
  • 项目类别:
    Priority Programmes
  • 资助金额:
    $0.0万
  • 财政年份:
    2017
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
FADeBaC Sentiment Analysis - Fully Automatic DEnsity-BAsed Clustering applied to Sentiment Analysis
  • 批准号:
    219327280
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2012
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
semisupervised coreference resolution
  • 批准号:
    104076539
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2009
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
WordGraph - Development of a unified graph-theoretical system for acquiring lexico-semantic phenomena
海外基金