Latent Part-of-Speech Sequences for Neural Machine Translation

Latent Part-of-Speech Sequences for Neural Machine Translation
复制标题

DOI:
10.18653/v1/d19-1072
复制
发表时间:
2019-08
期刊:
--
影响因子:
--
通讯作者:
Xuewen Yang;Yingru Liu;Dongliang Xie;Xin Wang;Niranjan Balasubramanian
Xuewen Yang;Yingru Liu;Dongliang Xie;Xin Wang;Niranjan Balasubramanian
中科院分区:
其他
文献类型:
--
作者:
Xuewen Yang;Yingru Liu;Dongliang Xie;Xin Wang;Niranjan Balasubramanian

文献摘要

被引文献

相似文献

学习目标语的句法结构可以提高神经机器翻译(NMT)的效率。然而,通过潜在变量结合句法增加了推理的复杂性,因为模型需要边缘化潜在的句法结构。为了避免这一点,模特们经常求助于贪婪的搜索,这只允许他们探索潜在空间的有限部分。在这项工作中,我们引入了一个新的潜在变量模型LaSyn,该模型捕捉了语法和语义之间的相互依赖,同时允许在潜在空间上进行有效和高效的推理。LaSyn去耦合了连续的潜在变量之间的直接依赖关系,这使得它的解码器能够穷尽地搜索潜在的句法选择,同时保持解码速度与潜在变量词汇的大小成正比。我们通过修改一个基于变换的NMT系统来实现LaSyn,并设计了一种神经期望最大化算法,该算法将词性信息作为潜在序列进行正则化。对四个不同的机器翻译任务的评估表明,将目标端句法与LaSyn相结合既提高了翻译质量,也提供了提高多样性的机会。
Learning target side syntactic structure has been shown to improve Neural Machine Translation (NMT). However, incorporating syntax through latent variables introduces additional complexity in inference, as the models need to marginalize over the latent syntactic structures. To avoid this, models often resort to greedy search which only allows them to explore a limited portion of the latent space. In this work, we introduce a new latent variable model, LaSyn, that captures the co-dependence between syntax and semantics, while allowing for effective and efficient inference over the latent space. LaSyn decouples direct dependence between successive latent variables, which allows its decoder to exhaustively search through the latent syntactic choices, while keeping decoding speed proportional to the size of the latent variable vocabulary. We implement LaSyn by modifying a transformer-based NMT system and design a neural expectation maximization algorithm that we regularize with part-of-speech information as the latent sequences. Evaluations on four different MT tasks show that incorporating target side syntax with LaSyn improves both translation quality, and also provides an opportunity to improve diversity.