Learning to Stop in Structured Prediction for Neural Machine Translation

Learning to Stop in Structured Prediction for Neural Machine Translation
复制标题

DOI:
10.18653/v1/n19-1187
复制
发表时间:
2019-04
期刊:
--
影响因子:
--
通讯作者:
Mingbo Ma;Renjie Zheng;Liang Huang
Mingbo Ma;Renjie Zheng;Liang Huang
中科院分区:
其他
文献类型:
--
作者:
Mingbo Ma;Renjie Zheng;Liang Huang

文献摘要

相似文献

束搜索优化(怀斯曼和拉什,2016年)解决了神经机器翻译中的许多问题。然而,这种方法缺乏有原则的停止标准,并且在训练过程中没有学习如何停止。实际上,在测试阶段,由于模型使用的是原始分数而非基于概率的分数,它自然会倾向于选择更长的假设。我们提出了一种新颖的排序方法,该方法能够实现最优的束搜索停止标准。我们进一步引入了一种结构化预测损失函数,该函数会对训练过程中束搜索产生的次优已完成候选结果进行惩罚。在合成数据以及真实语言(德译英和中译英)上进行的神经机器翻译实验表明,我们提出的方法在译文长度和BLEU分数方面都取得了更好的结果。
Beam search optimization (Wiseman and Rush, 2016) resolves many issues in neural machine translation. However, this method lacks principled stopping criteria and does not learn how to stop during training, and the model naturally prefers longer hypotheses during the testing time in practice since they use the raw score instead of the probability-based score. We propose a novel ranking method which enables an optimal beam search stop- ping criteria. We further introduce a structured prediction loss function which penalizes suboptimal finished candidates produced by beam search during training. Experiments of neural machine translation on both synthetic data and real languages (German→English and Chinese→English) demonstrate our pro- posed methods lead to better length and BLEU score.