Encoder-Decoder Attention ≠ Word Alignment: Axiomatic Method of Learning Word Alignments for Neural Machine Translation

Encoder-Decoder Attention ≠ Word Alignment: Axiomatic Method of Learning Word Alignments for Neural Machine Translation
复制标题

DOI:
10.5715/jnlp.27.531
复制
发表时间:
2020-09
期刊:
Journal of Natural Language Processing
影响因子:
--
通讯作者:
Chunpeng Ma;Akihiro Tamura;M. Utiyama;T. Zhao;E. Sumita
Chunpeng Ma;Akihiro Tamura;M. Utiyama;T. Zhao;E. Sumita
中科院分区:
其他
文献类型:
--
作者:
Chunpeng Ma;Akihiro Tamura;M. Utiyama;T. Zhao;E. Sumita

文献摘要

相似文献

在传统的神经机器翻译(NMT)模型中,如基于RNN的模型中,编解码器注意矩阵一直被视为(软)对齐模型。然而,我们的经验表明,这并不适用于变形金刚。通过将Transformer与基于RNN的NMT模型进行比较,我们发现了两个固有的差异,并相应地提出了两种捕获Transformer中字对齐的方法。此外,我们没有关注Transformer,而是提出了捕获词对齐的注意机制的三个公理,并基于这些公理提出了一种新的注意机制,我们称之为公理注意机制(AAM),它适用于任何NMT模型。AAM依赖于扰动函数,我们将几个扰动函数应用于AAM,包括基于掩码语言模型的新函数(Devlin,Chang,Lee和Toutanova 2019)。使用AAM指导自然机器翻译模型的训练,提高了自然机器翻译模型的翻译性能和词对齐的学习能力。我们的研究为神经机器翻译中序列到序列模型的解释提供了有益的启示。
The encoder-decoder attention matrix has been regarded as the (soft) alignment model for conventional neural machine translation (NMT) models such as RNN-based models. However, we show empirically that this is not true for the Transformer. By comparing the Transformer with the RNN-based NMT model, we find two inherent differences, and accordingly present two methods of capturing word alignments in the Transformer. Furthermore, instead of focusing on the Transformer, we present three axioms for the attention mechanism that captures word alignments, and propose a new attention mechanism based on these axioms that we have termed the axiomatic attention mechanism (AAM), and which is applicable to any NMT models. The AAM depends on a perturbation function, and we apply several perturbation functions to the AAM, including a novel function based on the masked language model (Devlin, Chang, Lee, and Toutanova 2019). Using the AAM to guide the training of an NMT model improved both the translation performance and the learning of word alignments of the NMT model. Our research sheds light on the interpretation of sequence-to-sequence models on neural machine translation.