Stronger Baselines for Trustable Results in Neural Machine Translation

Stronger Baselines for Trustable Results in Neural Machine Translation
复制标题

DOI:
10.18653/v1/w17-3203
复制
发表时间:
2017-06
影响因子:
19
通讯作者:
Michael J. Denkowski;Graham Neubig
Michael J. Denkowski;Graham Neubig
中科院分区:
材料科学1区
文献类型:
--
作者:
Michael J. Denkowski;Graham Neubig

文献摘要

被引文献

相似文献

随着神经机器翻译的有效性在语言和数据场景中得到证明,人们对神经机器翻译的兴趣迅速增长。新的研究定期引入架构和算法的改进,这些改进比“普通的”NMT实现带来了显著的收益。然而,这些新技术很少在以前发表的技术的背景下进行评估,特别是那些在最先进的生产和共享任务系统中广泛使用的技术。因此,通常很难确定研究的改进是否会延续到实际使用的系统中。在这项工作中,我们推荐了三种具体的方法,它们相对容易实施,并导致更强大的实验系统。除了报告显著提高的BLEU分数外,我们还深入分析了改进的来源以及基本NMT模型的固有弱点。然后,我们比较了文献中提出的几种其他技术在从香草系统开始与我们更强的基线时所提供的相对收益,表明实验结论可能会根据所选择的基线而改变。这表明选择一个强基线对于报告可靠的实验结果至关重要。
Interest in neural machine translation has grown rapidly as its effectiveness has been demonstrated across language and data scenarios. New research regularly introduces architectural and algorithmic improvements that lead to significant gains over “vanilla” NMT implementations. However, these new techniques are rarely evaluated in the context of previously published techniques, specifically those that are widely used in state-of-the-art production and shared-task systems. As a result, it is often difficult to determine whether improvements from research will carry over to systems deployed for real-world use. In this work, we recommend three specific methods that are relatively easy to implement and result in much stronger experimental systems. Beyond reporting significantly higher BLEU scores, we conduct an in-depth analysis of where improvements originate and what inherent weaknesses of basic NMT models are being addressed. We then compare the relative gains afforded by several other techniques proposed in the literature when starting with vanilla systems versus our stronger baselines, showing that experimental conclusions may change depending on the baseline chosen. This indicates that choosing a strong baseline is crucial for reporting reliable experimental results.