Statistical Significance Tests for Machine Translation Evaluation

Statistical Significance Tests for Machine Translation Evaluation
复制标题

DOI:
--
复制
发表时间:
2004-07
期刊:
--
影响因子:
--
通讯作者:
Philipp Koehn
Philipp Koehn
中科院分区:
其他
文献类型:
--
作者:
Philipp Koehn

文献摘要

被引文献

相似文献

如果两个翻译系统在测试集上的性能不同,我们能相信这表明真正的系统质量不同吗?为了回答这个问题,我们描述了自举响应方法来计算测试结果的统计显著性,并在BLEU分数的具体例子上验证它们。即使对于只有300个句子的小测试,我们的方法也可以保证测试结果的差异是真实的。
If two translation systems differ differ in performance on a test set, can we trust that this indicates a difference in true system quality? To answer this question, we describe bootstrap resampling methods to compute statistical significance of test results, and validate them on the concrete example of the BLEU score. Even for small test sizes of only 300 sentences, our methods may give us assurances that test result differences are real.