Synchronous Inference for Multilingual Neural Machine Translation

Synchronous Inference for Multilingual Neural Machine Translation
复制标题

多语言神经机器翻译的同步推理

DOI:
10.1109/taslp.2022.3178241
复制
发表时间:
2022
期刊:
IEEE/ACM transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Chengqing Zong
Chengqing Zong
中科院分区:
其他
文献类型:
--
作者:
Qian Wang;Chengqing Zong;Chengqing Zong

文献摘要

参考文献

相似文献

多语言神经机器翻译允许单个模型在多个语言对之间进行翻译,这大大降低了模型训练的成本,近年来受到了广泛的关注。以往的研究主要集中于训练阶段的优化和提高不同参数共享程度的语言之间的正向知识转移,而忽略了推理过程中的多语种知识转移,尽管一种语言的翻译可能有助于其他语言的生成。这项工作在推理阶段加强了多目标语言之间的知识共享。为了实现这一点,我们提出了一种同步推理方法,可以同时生成多种语言的翻译。在生成过程中,该模型不仅根据源句和先前预测的片段预测每种语言的下一个单词,而且还根据其他目标语言的预测单词进行预测。为了最大化推理阶段的知识共享,我们设计了跨语言注意模块,允许模型从多个目标语言中动态地选择最相关的信息。同步推理模型需要多路并行训练数据,这是一个稀缺的问题。因此,我们建议采用多任务学习来整合大规模的双语数据。我们在三个多语言翻译数据集上对我们的方法进行了测试,结果表明,与强双语和多语言基线相比,该方法显著提高了翻译质量和译码效率。
Multilingual neural machine translation allows a single model to translate between multiple language pairs, which greatly reduces the cost of model training and receives much attention recently. Previous studies mainly focus on training stage optimization and improve positive knowledge transfer among languages with different levels of parameter sharing, but ignore the multilingual knowledge transfer during inference although the translation in one language may help the generation of other languages. This work enhances knowledge sharing among multiple target languages in the inference phase. To achieve this, we propose a synchronous inference method that can simultaneously generate translations in multiple languages. During generation, the model predicts the next word of each language not only based on source sentence and previously predicted segments, but also based on predicted words of other target languages. To maximize the inference stage knowledge sharing, we design a cross-lingual attention module which allows the model to dynamically select the most relevant information from multiple target languages. The synchronous inference model requires multi-way parallel training data which is scarce. We therefore propose to adopt multi-task learning to incorporate large-scale bilingual data. We evaluate our method on three multilingual translation datasets and prove that the proposed method significantly improve the translation quality and the decoding efficiency compared to strong bilingual and multilingual baselines.
DOI: 10.1016/j.csl.2016.10.006
发表时间: 2017-09
期刊: Comput. Speech Lang.
影响因子: --
作者:
Orhan Firat;Kyunghyun Cho;B. Sankaran;F. Yarman-Vural;Yoshua Bengio
通讯作者: Orhan Firat;Kyunghyun Cho;B. Sankaran;F. Yarman-Vural;Yoshua Bengio
DOI: 10.18653/v1/d19-1330
发表时间: 2019-11
期刊: --
影响因子: --
作者:
Yining Wang;Jiajun Zhang;Long Zhou;Yuchen Liu-;Chengqing Zong
通讯作者: Yining Wang;Jiajun Zhang;Long Zhou;Yuchen Liu-;Chengqing Zong
DOI: 10.18653/v1/2020.emnlp-main.365
发表时间: 2020-04
期刊: ArXiv
影响因子: --
作者:
Nils Reimers;Iryna Gurevych
通讯作者: Nils Reimers;Iryna Gurevych
DOI: 10.18653/v1/n19-4009
发表时间: 2019-04
期刊: --
影响因子: --
作者:
Myle Ott;Sergey Edunov;Alexei Baevski;Angela Fan;Sam Gross;Nathan Ng;David Grangier;Michael Auli-Michael-Aul
通讯作者: Myle Ott;Sergey Edunov;Alexei Baevski;Angela Fan;Sam Gross;Nathan Ng;David Grangier;Michael Auli-Michael-Aul
神经机器翻译的过去和未来建模
DOI: 10.1162/tacl_a_00011
发表时间: 2017-11
期刊: TACL2018
影响因子: --
作者:
Zaixiang Zheng;Hao Zhou;Shujian Huang;Lili Mou;Xinyu Dai;Jiajun Chen;Zhaopeng Tu
通讯作者: Zhaopeng Tu