Combining Confidence Estimation and Reference-based Metrics for Segment-level MT Evaluation

Combining Confidence Estimation and Reference-based Metrics for Segment-level MT Evaluation
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Lucia Specia;J. Giménez
Lucia Specia;J. Giménez
中科院分区:
其他
文献类型:
--
作者:
Lucia Specia;J. Giménez

文献摘要

被引文献

相似文献

我们描述了一种改进机器翻译(MT)评估的标准参考指标的努力,方法是用置信度估计(CE)特征丰富它们,并使用基于人类注释训练的学习机制。基于参考的机器翻译评估指标将系统输出与参考翻译进行比较,寻找不同层次(词汇、句法和语义)上的重叠。这些指标旨在比较机器翻译系统或分析给定系统的进展,并且已知在语料库级别上与人类判断具有相当好的相关性,但在片段级别上则不然。另一方面,CE度量以使用中的系统为目标,为每个翻译段的最终用户提供质量分数。他们不能依赖于参考翻译,而是使用从输入文本、系统输出和可能的外部语料库中提取的信息来训练机器学习算法。这些指标与分段级别的人类判断有更好的关联。然而,它们通常因输入段的难度等级而高度偏差,因此不太适合比较翻译相同输入段的多个系统。我们表明,这两类指标是互补的,可以结合起来提供MT评估指标,在段水平上实现与人类判断的更高相关性。
We describe an effort to improve standard reference-based metrics for Machine Translation (MT) evaluation by enriching them with Confidence Estimation (CE) features and using a learning mechanism trained on human annotations. Reference-based MT evaluation metrics compare the system output against reference translations looking for overlaps at different levels (lexical, syntactic, and semantic). These metrics aim at comparing MT systems or analyzing the progress of a given system and are known to have reasonably good correlation with human judgments at the corpus level, but not at the segment level. CE metrics, on the other hand, target the system in use, providing a quality score to the end-user for each translated segment. They cannot rely on reference translations, and use instead information extracted from the input text, system output and possibly external corpora to train machine learning algorithms. These metrics correlate better with human judgments at the segment level. However, they are usually highly biased by difficulty level of the input segment, and therefore are less appropriate for comparing multiple systems translating the same input segments. We show that these two classes of metrics are complementary and can be combined to provide MT evaluation metrics that achieve higher correlation with human judgments at the segment level.