Word Error Rates: Decomposition over POS classes and Applications for Error Analysis

Word Error Rates: Decomposition over POS classes and Applications for Error Analysis
复制标题

DOI:
10.3115/1626355.1626362
复制
发表时间:
2007-06
期刊:
--
影响因子:
--
通讯作者:
Maja Popovic;H. Ney
Maja Popovic;H. Ney
中科院分区:
其他
文献类型:
--
作者:
Maja Popovic;H. Ney

文献摘要

被引文献

相似文献

机器翻译输出的评价和错误分析是一项重要而又困难的任务。在这项工作中,我们提出了一种新的方法,通过在不同的词性类别上分解词错误率(Wer)和位置无关的词错误率(PER)来获得更多关于生成输出中实际翻译错误的详细信息。此外,我们研究了使用这些分解进行自动错误分析的两个可能的方面:屈折错误的估计和遗漏单词在词性类别上的分布。所获得的结果与人为错误分析的结果相对应。在欧洲议会全体会议西班牙语和英语语料库上取得的结果更好地概述了翻译错误的性质,并提出了在哪里努力改进翻译系统的意见。
Evaluation and error analysis of machine translation output are important but difficult tasks. In this work, we propose a novel method for obtaining more details about actual translation errors in the generated output by introducing the decomposition of Word Error Rate (Wer) and Position independent word Error Rate (Per) over different Part-of-Speech (Pos) classes. Furthermore, we investigate two possible aspects of the use of these decompositions for automatic error analysis: estimation of inflectional errors and distribution of missing words over Pos classes. The obtained results are shown to correspond to the results of a human error analysis. The results obtained on the European Parliament Plenary Session corpus in Spanish and English give a better overview of the nature of translation errors as well as ideas of where to put efforts for possible improvements of the translation system.