Extremely low-resource neural machine translation for Asian languages

Extremely low-resource neural machine translation for Asian languages
复制标题

DOI:
10.1007/s10590-020-09258-6
复制
发表时间:
2021-02-10
影响因子:
1.9
通讯作者:
Sumita, Eiichiro
Sumita, Eiichiro
中科院分区:
其他
文献类型:
--
作者:
Rubino, Raphael;Marie, Benjamin;Sumita, Eiichiro

文献摘要

被引文献

相似文献

本文提出了一套有效的方法来处理极低资源的语言对的自注意神经机器翻译(NMT)专注于英语和亚洲四种语言。从最初用于训练双语基线模型的一组平行句子开始,我们引入了额外的单语语料库和数据处理技术来提高翻译质量。我们描述了一系列最佳实践,并通过对八个翻译方向进行评估,根据最先进的NMT方法(如超参数搜索,结合标签和噪声的前向和后向翻译的数据增强,以及联合多语言培训)对方法进行了经验验证。实验表明,常用的自注意NMT模型的默认架构并没有达到最佳的结果,验证了以前的工作的超参数调整的重要性。此外,经验结果表明,需要大量的合成数据来有效地增加模型的参数,从而获得由自动度量衡量的最佳翻译质量。我们发现,在大量标记的反向翻译上训练的最佳NMT模型优于其他三种合成数据生成方法。最后,与统计机器翻译(SMT)的比较表明,极低的资源NMT需要大量的合成并行数据与反向翻译获得,以缩小性能差距与前面的SMT方法。
This paper presents a set of effective approaches to handle extremely low-resource language pairs for self-attention based neural machine translation (NMT) focusing on English and four Asian languages. Starting from an initial set of parallel sentences used to train bilingual baseline models, we introduce additional monolingual corpora and data processing techniques to improve translation quality. We describe a series of best practices and empirically validate the methods through an evaluation conducted on eight translation directions, based on state-of-the-art NMT approaches such as hyper-parameter search, data augmentation with forward and backward translation in combination with tags and noise, as well as joint multilingual training. Experiments show that the commonly used default architecture of self-attention NMT models does not reach the best results, validating previous work on the importance of hyper-parameter tuning. Additionally, empirical results indicate the amount of synthetic data required to efficiently increase the parameters of the models leading to the best translation quality measured by automatic metrics. We show that the best NMT models trained on large amount of tagged back-translations outperform three other synthetic data generation approaches. Finally, comparison with statistical machine translation (SMT) indicates that extremely low-resource NMT requires a large amount of synthetic parallel data obtained with back-translation in order to close the performance gap with the preceding SMT approach.