Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers

Scientific Credibility of Machine Translation Research: A Meta-Evaluation of 769 Papers
复制标题

DOI:
10.18653/v1/2021.acl-long.566
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Benjamin Marie;Atsushi Fujita;Raphaël Rubino
Benjamin Marie;Atsushi Fujita;Raphaël Rubino
中科院分区:
其他
文献类型:
--
作者:
Benjamin Marie;Atsushi Fujita;Raphaël Rubino

文献摘要

相似文献

本文介绍了第一次大规模的机器翻译元评估。我们对2010年至2020年发表的769篇研究论文进行了MT评价。我们的研究表明,在过去的十年中,自动MT评估的实践发生了巨大的变化,并遵循有关的趋势。越来越多的MT评估完全依赖于BLEU分数之间的差异来得出结论,而不进行任何统计显著性测试或人工评估,同时至少有108个指标声称比BLEU更好。在最近的论文中,MT评估倾向于复制和比较以前工作中的自动度量分数,以声称方法或算法的优越性,而不确认是否使用了完全相同的训练,验证和测试数据,也没有度量分数是可比的。此外,报告标准化指标分数的工具还远未被机器翻译界广泛采用。在展示了这些陷阱的积累如何导致可疑的评估之后,我们提出了一个指导方针,以鼓励更好的自动MT评估沿着一个简单的元评估评分方法来评估其可信度。
This paper presents the first large-scale meta-evaluation of machine translation (MT). We annotated MT evaluations conducted in 769 research papers published from 2010 to 2020. Our study shows that practices for automatic MT evaluation have dramatically changed during the past decade and follow concerning trends. An increasing number of MT evaluations exclusively rely on differences between BLEU scores to draw conclusions, without performing any kind of statistical significance testing nor human evaluation, while at least 108 metrics claiming to be better than BLEU have been proposed. MT evaluations in recent papers tend to copy and compare automatic metric scores from previous work to claim the superiority of a method or an algorithm without confirming neither exactly the same training, validating, and testing data have been used nor the metric scores are comparable. Furthermore, tools for reporting standardized metric scores are still far from being widely adopted by the MT community. After showing how the accumulation of these pitfalls leads to dubious evaluation, we propose a guideline to encourage better automatic MT evaluation along with a simple meta-evaluation scoring method to assess its credibility.