Syntax-based Transformer for Neural Machine Translation

Syntax-based Transformer for Neural Machine Translation
复制标题

DOI:
10.5715/jnlp.27.445
复制
发表时间:
2020-06
期刊:
Journal of Natural Language Processing
影响因子:
--
通讯作者:
Chunpeng Ma;Akihiro Tamura;M. Utiyama;E. Sumita;T. Zhao
Chunpeng Ma;Akihiro Tamura;M. Utiyama;E. Sumita;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Chunpeng Ma;Akihiro Tamura;M. Utiyama;E. Sumita;T. Zhao

文献摘要

相似文献

Transformer(Vaswani,Shazeer,Parmar,Uszkoreit,Jones,Gomez,Kaiser,and Polosukhin 2017)完全依赖于注意力机制,在机器翻译(MT)上实现了最先进的性能。然而,语法信息,这已经改善了许多以前的MT模型,并没有明确利用Transformer。我们提出了一个基于语法的Transformer的MT,它将源端的语法结构生成的解析器到编码器的自我注意和位置编码。我们的方法是通用的,因为它适用于组成树和包装森林。两个语言对的评估表明,我们的基于语法的Transformer优于传统的(非语法)Transformer。BLEUs在英日、英中和英德翻译任务上的改进分别达到2.32、2.91和1.03。进一步的消融研究和定性分析表明,基于句法的自我注意在学习局部结构信息方面表现较好,而基于句法的位置编码在学习全局结构信息方面表现较好。
The Transformer (Vaswani, Shazeer, Parmar, Uszkoreit, Jones, Gomez, Kaiser, and Polosukhin 2017), which purely depends on attention mechanism, has achieved stateof-the-art performance on machine translation (MT). However, syntactic information, which has improved many previous MT models, has not been utilized explicitly by Transformer. We propose a syntax-based Transformer for MT, which incorporates source-side syntax structures generated by the parser into the self-attention and positional encoding of the encoder. Our method is general in that it is applicable to both constituent trees and packed forests. Evaluations on two language pairs show that our syntax-based Transformer outperforms the conventional (non-syntactic) Transformer. The improvements of BLEUs on English-Japanese, English-Chinese and English-German translation tasks are up to 2.32, 2.91 and 1.03, respectively. Furthermore, our ablation study and qualitative analysis demonstrate that the syntax-based self-attention does well in learning local structural information, while the syntax-based positional encoding does well in learning global structural information.