A STUDY OF TRANSLATION ERROR RATE WITH TARGETED HUMAN ANNOTATION

A STUDY OF TRANSLATION ERROR RATE WITH TARGETED HUMAN ANNOTATION
复制标题

DOI:
--
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
M. Snover;B. Dorr;R. Schwartz;Linnea Micciulla;R. Weischedel
M. Snover;B. Dorr;R. Schwartz;Linnea Micciulla;R. Weischedel
中科院分区:
其他
文献类型:
--
作者:
M. Snover;B. Dorr;R. Schwartz;Linnea Micciulla;R. Weischedel

文献摘要

被引文献

相似文献

我们定义了一种新的、直观的衡量标准来评估机器翻译输出,避免了更多基于意义的方法的知识密集型和人类判断的劳动密集型。翻译错误率(TER)衡量的是人类为了改变系统输出,使其与参考翻译完全匹配而必须执行的编辑量。我们还计算人类目标TER(或HTER),其中针对人类“目标参考”计算翻译的最小TER,该参考保留含义(由参考翻译提供)并且是流畅的,但是被选择以最小化特定系统输出的TER分数。我们证明:(1)TER的单参考变体与BLEU的四参考变体一样与MT质量的人类判断相关;(2)人类靶向HTER产生33%的错误率降低,并且被证明与人类判断非常好地相关;(3)TER的四参考变体和HTER的单参考变体与人类判断的相关性高于BLEU;(4)HTER与人类判断的相关性比METEOR或其人类靶向变体(HMETEOR)更高;(5)TER的四参考变体与第二人类判断的单一人类判断相关,而HTER,HBLEU和HMETEOR与人类判断的相关性明显优于第二人类判断。
We define a new, intuitive measure for evaluating machine translation output that avoids the knowledge intensiveness of more meaning-based approaches, and the labor-intensiveness of human judgments. Translation Error Rate (TER) measures the amount of editing that a human would have to perform to change a system output so it exactly matches a reference translation. We also compute a human-targeted TER (or HTER), where the minimum TER of the translation is computed against a human ‘targeted reference’ that preserves the meaning (provided by the reference translations) and is fluent, but is chosen to minimize the TER score for a particular system output. We show that: (1) The single-reference variant of TER correlates as well with human judgments of MT quality as the four-reference variant of BLEU; (2) The human-targeted HTER yields a 33% error-rate reduction and is shown to be very well correlated with human judgments; (3) The four-reference variant of TER and the single-reference variant of HTER yield higher correlations with human judgments than BLEU; (4) HTER yields higher correlations with human judgments than METEOR or its human-targeted variant (HMETEOR); and (5) The four-reference variant of TER correlates as well with a single human judgment as a second human judgment does, while HTER, HBLEU, and HMETEOR correlate significantly better with a human judgment than a second human judgment does.