Dependency-Based Self-Attention for Transformer NMT

Dependency-Based Self-Attention for Transformer NMT
复制标题

DOI:
10.26615/978-954-452-056-4_028
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Hiroyuki Deguchi;Akihiro Tamura;Takashi Ninomiya
Hiroyuki Deguchi;Akihiro Tamura;Takashi Ninomiya
中科院分区:
其他
文献类型:
--
作者:
Hiroyuki Deguchi;Akihiro Tamura;Takashi Ninomiya

文献摘要

相似文献

在本文中,我们提出了一种新的Transformer神经机器翻译(NMT)模型,该模型将依赖关系纳入源端和目标端的自注意中,即基于依赖的自注意。基于依赖的自注意受LISA (linguistic - informed self-attention)的启发,在依赖关系的约束下被训练去关注每个标记的被修饰语。LISA最初是为Transformer编码器的语义角色标注而提出的,本文将LISA扩展到Transformer NMT,通过在解码器端基于依赖的自关注中屏蔽单词的未来信息。此外,我们基于依赖的自我关注在由字节对编码创建的子字单元上运行。实验表明,我们的模型在WAT’18亚洲科学论文摘录语料库的日语到英语翻译任务上比基线模型提高了1.0个BLEU点。
In this paper, we propose a new Transformer neural machine translation (NMT) model that incorporates dependency relations into self-attention on both source and target sides, dependency-based self-attention. The dependency-based self-attention is trained to attend to the modifiee for each token under constraints based on the dependency relations, inspired by Linguistically-Informed Self-Attention (LISA). While LISA is originally proposed for Transformer encoder for semantic role labeling, this paper extends LISA to Transformer NMT by masking future information on words in the decoder-side dependency-based self-attention. Additionally, our dependency-based self-attention operates at sub-word units created by byte pair encoding. The experiments show that our model improves 1.0 BLEU points over the baseline model on the WAT’18 Asian Scientific Paper Excerpt Corpus Japanese-to-English translation task.