Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation

Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation
复制标题

DOI:
10.1145/3491065
复制
发表时间:
2022-01
期刊:
Transactions on Asian and Low-Resource Language Information Processing
影响因子:
--
通讯作者:
Zhuoyuan Mao;Chenhui Chu;S. Kurohashi
Zhuoyuan Mao;Chenhui Chu;S. Kurohashi
中科院分区:
其他
文献类型:
--
作者:
Zhuoyuan Mao;Chenhui Chu;S. Kurohashi

文献摘要

相似文献

在本研究中,我们提出了针对低资源机器翻译(NMT)的新的序列到序列的预训练目标:针对以日语为源语言或目标语言的语言对的日语特定的序列到序列(JASS),以及针对涉及英语的语言对的英语特定的序列到序列(ENSS)。JASS侧重于掩蔽和重新排序日语语言单位,而ENSS则是基于短语结构掩蔽和重新排序任务而提出的。在ASPEC日语-英语和日语-汉语、维基百科日语-汉语和新闻英语-韩语语料库上的实验表明,JASS和ENSS在日语-英语任务上优于MASS和其他语言不可知的预训练方法,在日语-英语任务上高达+2.9点,在日语-汉语任务上高达+7.0BLEU分,在英-韩语任务上高达+1.3点BLEU。实证分析侧重于JASS和ENSS中各个部分之间的关系,揭示了JASS和ENSS子任务的互补性。使用LASeR、人工评估和案例研究的充分性评估表明,我们提出的方法显著优于没有注入语言知识的预训练方法,并且与流利性相比,它们对充分性有更大的积极影响。
In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese–English & Japanese–Chinese, Wikipedia Japanese–Chinese, News English–Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese–English tasks, up to +7.0 BLEU points for the Japanese–Chinese tasks and up to +1.3 BLEU points for English–Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency.