Hard Non-Monotonic Attention for Character-Level Transduction

Hard Non-Monotonic Attention for Character-Level Transduction
复制标题

字符级转换的硬非单调注意力

DOI:
--
复制
发表时间:
2018
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Ryan Cotterell
Ryan Cotterell
中科院分区:
--
文献类型:
--
作者:
Shijie Wu;Pamela Shapiro;Ryan Cotterell

文献摘要

被引文献

相似文献

字符级字符串到字符串的转换是各种自然语言处理任务的重要组成部分。目标是将输入字符串映射到输出字符串,其中字符串可能具有不同的长度,并且具有来自不同字母的字符。最近的方法使用带有注意机制的序列到序列模型来学习模型在生成输出字符串时应该关注输入字符串的哪些部分。软注意和硬单调注意都已被使用,但硬非单调注意仅用于其他序列建模任务,并且需要随机逼近来计算梯度。在这项工作中,我们引入了一种精确的多项式时间算法来边缘化两个字符串之间的非单调排列的指数数量,表明硬注意模型可以被视为经典IBM模型1的神经重参数化。我们通过实验比较了软非单调注意和硬非单调注意,发现精确算法比随机近似显著提高了性能,并且优于软注意。
Character-level string-to-string transduction is an important component of various NLP tasks. The goal is to map an input string to an output string, where the strings may be of different lengths and have characters taken from different alphabets. Recent approaches have used sequence-to-sequence models with an attention mechanism to learn which parts of the input string the model should focus on during the generation of the output string. Both soft attention and hard monotonic attention have been used, but hard non-monotonic attention has only been used in other sequence modeling tasks and has required a stochastic approximation to compute the gradient. In this work, we introduce an exact, polynomial-time algorithm for marginalizing over the exponential number of non-monotonic alignments between two strings, showing that hard attention models can be viewed as neural reparameterizations of the classical IBM Model 1. We compare soft and hard non-monotonic attention experimentally and find that the exact algorithm significantly improves performance over the stochastic approximation and outperforms soft attention.