课题基金 / 基金详情

Models of morphosyntax for statistical machine translation

Models of morphosyntax for statistical machine translation
统计机器翻译的形态句法模型
批准号:
123083856
负责人:
Professor Dr. Hinrich Schütze
金额:
$0.0万
依托单位国家:
德国
项目类别:
Research Grants
财政年份:
2009
资助国家:
德国
项目状态:
已结题
起止时间:
2008-12-31 至 2017-12-31

项目摘要

项目成果

Professor Dr. Hinrich Schütze的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Statistical approaches to machine translation (MT) have shownthemselves to be effective in the last few years. However, whentranslating into a morphologically rich language this is not true,particularly when there is also significant syntactic divergencebetween the two languages. The quality of statistical machinetranslation (SMT) is poor in this case because of independenceassumptions made between the models of morphology, syntax andtranslation that do not reflect linguistic reality.In the first phase of the project we made significant strides intranslating into German, a morphologically rich language. We focusedon issues of linguistic representation and linguistic resources withinstatistical machine translation. We carried out original research indealing with German word formation (addressing both compounds andportmanteaus), German inflectional morphology and syntactic issues indealing with both English to German translation and German to Englishtranslation. We published seven conference publications at top rankedinternational conferences as well as two workshop contributions, andalso supervised Bachelors-level and Masters-level student workrelevant to the project.In the proposed phase 2, we will move on from afocus on linguistic representation to working on advanced machinelearning approaches for solving the linguistic problems inherent inthe difficult machine translation language pair English/German. In theprevious phase of the work, we focused on general linguistic problemsin translation. One of the most important lessons we learned inanalyzing the output of our linguistically enhanced systems is thatthe issue of the mismatch of domains between training data and testdata is a critically important issue. The training data is mostlytaken from the European parliament proceedings, but the testing datais from the news domain or many other domains (including the medicaldomain, which we will study in this phase of the project).In the proposed phase 2 of the Morphosyntax project, we will followfour main lines of work. We will extend our successful work on Germanword formation and inflectional morphology by reducing our dependenceon hand-crafted morphological resources by determining how to performsemi-supervised acquisition of morphological resources, payingparticular attention to the important issue of domainadaptation. Within hierarchical decoding (wheresyntactic formalisms are used as the representation for translation),we will study the integration of advanced machine learning methods forchoosing syntactic reorderings. We will also study the issue of addingsemantic information (in addition to syntactic information) intohierarchical decoding. Finally, we will study different ways tointegrate powerful classification approaches (required for the otherwork packages) directly into the decoder rather than using externalpre-processing or post-processing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ReMLAV: Relational Machine Learning for Argument Validation
  • 批准号:
    376183703
  • 项目类别:
    Priority Programmes
  • 资助金额:
    $0.0万
  • 财政年份:
    2017
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
FADeBaC Sentiment Analysis - Fully Automatic DEnsity-BAsed Clustering applied to Sentiment Analysis
  • 批准号:
    219327280
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2012
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
semisupervised coreference resolution
  • 批准号:
    104076539
  • 项目类别:
    Research Grants
  • 资助金额:
    $0.0万
  • 财政年份:
    2009
  • 负责人:
    Professor Dr. Hinrich Schütze
  • 依托单位:
WordGraph - Development of a unified graph-theoretical system for acquiring lexico-semantic phenomena
海外基金