课题基金 / 基金详情

Discriminative Phrase-Based Statistical Machine Translation

Discriminative Phrase-Based Statistical Machine Translation
基于判别性短语的统计机器翻译
批准号:
EP/D074959/1
负责人:
M Osborne
金额:
$34.31万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2007
资助国家:
英国
项目状态:
已结题
起止时间:
2007 至 --

项目摘要

项目成果

M Osborne的其他基金

相似基金

相关文献

中文摘要
翻译
统计机器翻译(SMT)在过去十年中取得了很大的进步。这些系统的一个显著特点是它们很少使用关于翻译的语言学知识。所有关于如何翻译句子的知识都以数据驱动的方式从平行语料库(句子与其翻译配对)中收集。根据这一观察,我们可以看到,可用于训练的并行语料库的数量不会以实质性的速度增加。这表明,SMT的进一步发展将来自对现有数据的更好建模:这意味着将语言学引入翻译问题。对于某些语言,很容易获得语言约束。对于其他语言,这种信息不太普遍。我们想看看,即使使用贫乏的知识来源,翻译是否也能取得进步。为了成功地执行这种集成,我们需要一个灵活的框架。我们将使用判别机器学习技术(最大熵)扩展现有方法(产生最先进的结果)。这些方法不仅可以让我们轻松地将语言学融入翻译过程,而且还可以让我们通过更好的建模来改进最先进的技术。与更好的建模相关的是严重的缩放问题,我们有处理这些问题的经验。我们将研究的语言对包括德语-英语、阿拉伯语-英语和汉语-英语。最后,我们将参加国际机器翻译评估练习。这将涉及对我们的翻译质量进行自动和手动评估,并将我们的方法与其他小组的方法进行比较。
英文摘要
Statistical Machine Translation (SMT) has made great improvements over the last decade. A striking property of these systems is that they make minimal usage of linguistic knowledge about translation. All knowledge about how to translate sentences is gathered in a data-driven manner from parallel corpora (sentences paired with their translation).In tandem with this observation, projecting ahead, we can see that the volumes of parallel corpora available for traning will not increase at a substantial rate. This suggests that further progress in SMT will come from better modelling of the existing data we have: this means bringing linguistics to the translation problem.For some languages, linguistic constraints are easily obtained. For other languages, this information is less widely present. We intend seeing whether an improvement in translation can be obtained even when using impoverished knowledge sources.To successfully carry out this integration, we need a flexible framework. We shall extend an existing approach (which yields state-of-the-art results) using techniques from discriminitive machine learning techinques ( maximim entropy ). These approaches will not only allow us to easily integrate linguistics into the translation process, but should also allow us to improve upon the state-of-the-art simply from better modelling. Associated with better modelling are serious scaling problems, for which we have experience at tackling.The language pairs we shall investigate will include German-English, Arabic-English and Chinese-English.Finally, we shall compete in international Machine Translation evaluation exercises. This will involve automatic and manual evaluation of our translation quality, and will allow comparison of our approaches with that of other groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ReDites: Real Time, Detection, Tracking, Monitoring and Interpretation of Events in Social Media
  • 批准号:
    EP/L010690/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $30.81万
  • 财政年份:
    2013
  • 负责人:
    M Osborne
  • 依托单位:
CROSS: Real-time Story Detection Across Multiple Massive Streams
  • 批准号:
    EP/J020664/1
  • 项目类别:
    Research Grant
  • 资助金额:
    $26.72万
  • 财政年份:
    2012
  • 负责人:
    M Osborne
  • 依托单位:
海外基金