A Study of Translation Edit Rate with Targeted Human Annotation

A Study of Translation Edit Rate with Targeted Human Annotation
复制标题

DOI:
--
复制
发表时间:
2006
期刊:
--
影响因子:
--
通讯作者:
M. Snover;B. Dorr;Richard M. Schwartz;Linnea Micciulla;J. Makhoul
M. Snover;B. Dorr;Richard M. Schwartz;Linnea Micciulla;J. Makhoul
中科院分区:
其他
文献类型:
--
作者:
M. Snover;B. Dorr;Richard M. Schwartz;Linnea Micciulla;J. Makhoul

文献摘要

被引文献

相似文献

我们研究了一种新的、直观的评估机器翻译输出的方法,它避免了更多基于意义的方法的知识密集度,以及人类判断的劳动密集度。翻译编辑率(TER)衡量人类为更改系统输出以使其与参考翻译完全匹配而必须执行的编辑量。我们发现,作为BLEU的四个参考变量,TER的单参考变量也与人类对MT质量的判断相关。我们还定义了以人类为目标的TER(或HTER),并表明它与人类判断的相关性高于BLEU--即使BLEU被给予以人类为目标的参考。我们的结果表明,HTER与人类判断的相关性比HMETEOR更好,TER和HTER的四个参考变体与人类判断的相关性与第二个人类判断一样-或者更好。
We examine a new, intuitive measure for evaluating machine-translation output that avoids the knowledge intensiveness of more meaning-based approaches, and the labor-intensiveness of human judgments. Translation Edit Rate (TER) measures the amount of editing that a human would have to perform to change a system output so it exactly matches a reference translation. We show that the single-reference variant of TER correlates as well with human judgments of MT quality as the four-reference variant of BLEU. We also define a human-targeted TER (or HTER) and show that it yields higher correlations with human judgments than BLEU—even when BLEU is given human-targeted references. Our results indicate that HTER correlates with human judgments better than HMETEOR and that the four-reference variants of TER and HTER correlate with human judgments as well as—or better than—a second human judgment does.