Parallel Attention Mechanisms in Neural Machine Translation

Parallel Attention Mechanisms in Neural Machine Translation
复制标题

DOI:
10.1109/icmla.2018.00088
复制
发表时间:
2018-10
期刊:
2018 17th IEEE International Conference on Machine Learning and Applications (ICMLA)
影响因子:
--
通讯作者:
Julian R. Medina;J. Kalita
Julian R. Medina;J. Kalita
中科院分区:
其他
文献类型:
--
作者:
Julian R. Medina;J. Kalita

文献摘要

被引文献

相似文献

最近的神经机器翻译论文提出,相对于先前的标准,例如循环神经网络和卷积神经网络(RNN 和 CNN),严格使用注意力机制。我们建议通过传统方式运行?并行堆叠来自编码器-解码器注意力集中架构的编码分支,可以从模型中删除更多顺序操作,从而减少训练时间。特别是,我们修改了 Google 最近发布的名为 Transformer 的基于注意力的架构,通过用并行注意力模块替换顺序注意力模块,减少了训练时间,同时大幅提高了 BLEU 分数。对英语到德语和英语到法语翻译任务的实验表明,我们的模型建立了新的技术水平。
Recent papers in neural machine translation have proposed the strict use of attention mechanisms over previous standards such as recurrent and convolutional neural networks (RNNs and CNNs). We propose that by running traditionally? stacked encoding branches from encoder-decoder attention-focused architectures in parallel, that even more sequential operations can be removed from the model and thereby decrease training time. In particular, we modify the recently published attention-based architecture called Transformer by Google, by replacing sequential attention modules with parallel ones, reducing the amount of training time and substantially improving BLEU scores at the same time. Experiments over the English to German and English to French translation tasks show that our model establishes a new state of the art.