Fluency, Adequacy, or HTER? Exploring Different Human Judgments with a Tunable MT Metric

Fluency, Adequacy, or HTER? Exploring Different Human Judgments with a Tunable MT Metric
复制标题

DOI:
10.3115/1626431.1626480
复制
发表时间:
2009-03
期刊:
--
影响因子:
--
通讯作者:
M. Snover;Nitin Madnani;B. Dorr;R. Schwartz
M. Snover;Nitin Madnani;B. Dorr;R. Schwartz
中科院分区:
其他
文献类型:
--
作者:
M. Snover;Nitin Madnani;B. Dorr;R. Schwartz

文献摘要

被引文献

相似文献

传统上,自动机器翻译(MT)的评价指标是通过给机器翻译输出的分数与人类对翻译性能的判断的相关性来评估的。不同类型的人类判断,如流利性、充分性和HTER,衡量机器翻译性能的不同方面,这些方面可以通过自动机器翻译度量来捕获。我们通过使用一种新的可调MT度量来探索这些差异:TER-Plus,它扩展了翻译编辑率评估度量,具有可调参数,并结合了形态学,同义词和释义。TER-Plus被证明是NIST的metrics MATR 2008挑战赛中的顶级指标之一,在Pearson和Spearman相关性方面具有最高的平均排名。将TER-Plus优化到不同类型的人类判断,可以显著提高不同类型编辑的相关性,并产生有意义的权重变化,表明人类判断类型之间存在显著差异。
Automatic Machine Translation (MT) evaluation metrics have traditionally been evaluated by the correlation of the scores they assign to MT output with human judgments of translation performance. Different types of human judgments, such as Fluency, Adequacy, and HTER, measure varying aspects of MT performance that can be captured by automatic MT metrics. We explore these differences through the use of a new tunable MT metric: TER-Plus, which extends the Translation Edit Rate evaluation metric with tunable parameters and the incorporation of morphology, synonymy and paraphrases. TER-Plus was shown to be one of the top metrics in NIST's Metrics MATR 2008 Challenge, having the highest average rank in terms of Pearson and Spearman correlation. Optimizing TER-Plus to different types of human judgments yields significantly improved correlations and meaningful changes in the weight of different types of edits, demonstrating significant differences between the types of human judgments.