Target Conditioned Sampling: Optimizing Data Selection for Multilingual Neural Machine Translation

Target Conditioned Sampling: Optimizing Data Selection for Multilingual Neural Machine Translation
复制标题

目标条件采样:优化多语言神经机器翻译的数据选择

DOI:
10.18653/v1/p19-1583
复制
发表时间:
2019
影响因子:
1.1
通讯作者:
Graham Neubig
Graham Neubig
中科院分区:
人文科学4区
文献类型:
--
作者:
Xinyi Wang;Graham Neubig

文献摘要

被引文献

相似文献

为了改进具有多语言语料库的低资源神经机器翻译(NMT),仅对最相关的高资源语言进行训练通常比使用所有可用数据更有效(Neubig和Hu,2018)。然而,它仍然是一个问题,一个智能的数据选择策略是否可以进一步改善低资源NMT与其他辅助语言的数据。在本文中,我们试图在所有多语言数据上构建一个抽样分布,以便最大限度地减少低资源语言的训练损失。在此基础上,我们提出了一种有效的算法(TCS),它首先对目标句子进行采样,然后对其源句子进行条件采样。实验表明,TCS在我们测试的四种语言中的三种语言上带来了高达2个BLEU改进的显着收益,并且训练开销最小。
To improve low-resource Neural Machine Translation (NMT) with multilingual corpus, training on the most related high-resource language only is generally more effective than us- ing all data available (Neubig and Hu, 2018). However, it remains a question whether a smart data selection strategy can further improve low-resource NMT with data from other auxiliary languages. In this paper, we seek to construct a sampling distribution over all multilingual data, so that it minimizes the training loss of the low-resource language. Based on this formulation, we propose and efficient algorithm, (TCS), which first samples a target sentence, and then conditionally samples its source sentence. Experiments show TCS brings significant gains of up to 2 BLEU improvements on three of four languages we test, with minimal training overhead.