Interrogating the Explanatory Power of Attention in Neural Machine Translation

Interrogating the Explanatory Power of Attention in Neural Machine Translation
复制标题

质疑神经机器翻译中注意力的解释力

DOI:
--
复制
发表时间:
2019
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Anoop Sarkar
Anoop Sarkar
中科院分区:
--
文献类型:
--
作者:
Pooya Moradi;Nishant Kambhatla;Anoop Sarkar

文献摘要

参考文献

被引文献

相似文献

注意力模型已成为神经机器翻译(NMT)的重要组成部分。它们通常隐式或显式地用于证明模型生成特定标记的决策的合理性,但尚未严格确定注意力在 NMT 中的可靠信息来源的程度。为了评估 NMT 注意力的解释力,我们研究了产生相同预测的可能性,但使用反事实注意力模型修改了训练后的注意力模型的关键方面。使用这些反事实注意机制,我们评估它们在翻译过程中仍然保留功能词和实词生成的程度。与最先进的注意力模型相比,我们的反事实注意力模型在我们的德语-英语数据集中产生了 68% 的虚词和 21% 的实词。我们的实验表明,注意力模型本身无法可靠地解释 NMT 模型做出的决策。
Attention models have become a crucial component in neural machine translation (NMT). They are often implicitly or explicitly used to justify the model’s decision in generating a specific token but it has not yet been rigorously established to what extent attention is a reliable source of information in NMT. To evaluate the explanatory power of attention for NMT, we examine the possibility of yielding the same prediction but with counterfactual attention models that modify crucial aspects of the trained attention model. Using these counterfactual attention mechanisms we assess the extent to which they still preserve the generation of function and content words in the translation process. Compared to a state of the art attention model, our counterfactual attention models produce 68% of function words and 21% of content words in our German-English dataset. Our experiments demonstrate that attention models by themselves cannot reliably explain the decisions made by a NMT model.
DOI: 10.1016/j.patcog.2016.11.008
发表时间: 2017-05-01
影响因子: 8
作者:
Montavon, Gregoire;Lapuschkin, Sebastian;Mueller, Klaus-Robert
通讯作者: Mueller, Klaus-Robert