Latent Alignment and Variational Attention

Latent Alignment and Variational Attention
复制标题

DOI:
--
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
Yuntian Deng;Yoon Kim;Justin T Chiu;Demi Guo;Alexander M. Rush
Yuntian Deng;Yoon Kim;Justin T Chiu;Demi Guo;Alexander M. Rush
中科院分区:
其他
文献类型:
--
作者:
Yuntian Deng;Yoon Kim;Justin T Chiu;Demi Guo;Alexander M. Rush

文献摘要

被引文献

相似文献

神经注意力已经成为自然语言处理和相关领域许多最先进模型的核心。注意力网络是一种易于训练且有效的软模拟对齐方法;然而,该方法并没有在概率意义上边缘化潜在的对齐。这一特性使得将注意力与其他对齐方法进行比较、将其与概率模型组合以及根据观察到的数据执行后验推理变得困难。相关的潜在方法,即硬注意力,可以解决这些问题,但通常更难训练且不太准确。这项工作考虑了变分注意力网络,它是学习潜在变量对齐模型的软注意力和硬注意力的替代品,具有基于摊销变分推理的更严格的近似界限。我们进一步提出了减少梯度方差的方法,以使这些方法在计算上可行。实验表明,对于机器翻译和视觉问答,低效的精确潜变量模型优于标准神经注意力,但当使用基于硬注意力的训练时,这些收益就会消失。另一方面,变分注意力保留了大部分性能增益,但训练速度与神经注意力相当。
Neural attention has become central to many state-of-the-art models in natural language processing and related domains. Attention networks are an easy-to-train and effective method for softly simulating alignment; however, the approach does not marginalize over latent alignments in a probabilistic sense. This property makes it difficult to compare attention to other alignment approaches, to compose it with probabilistic models, and to perform posterior inference conditioned on observed data. A related latent approach, hard attention, fixes these issues, but is generally harder to train and less accurate. This work considers variational attention networks, alternatives to soft and hard attention for learning latent variable alignment models, with tighter approximation bounds based on amortized variational inference. We further propose methods for reducing the variance of gradients to make these approaches computationally feasible. Experiments show that for machine translation and visual question answering, inefficient exact latent variable models outperform standard neural attention, but these gains go away when using hard attention based training. On the other hand, variational attention retains most of the performance gain but with training speed comparable to neural attention.