Making attention mechanisms more robust and interpretable with virtual adversarial training

Making attention mechanisms more robust and interpretable with virtual adversarial training
复制标题

DOI:
10.1007/s10489-022-04301-w
复制
发表时间:
2021-04
影响因子:
5.3
通讯作者:
Shunsuke Kitada;H. Iyatomi
Shunsuke Kitada;H. Iyatomi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Shunsuke Kitada;H. Iyatomi

文献摘要

相似文献

虽然注意力机制已经成为深度学习模型的基本组成部分,但它们容易受到干扰,这可能会降低预测性能和模型的可解释性。注意力机制的对抗训练(AT)通过考虑对抗性扰动成功地减少了这些缺点。然而,这种技术需要标签信息,因此,其使用仅限于监督设置。在这项研究中,我们探讨了将虚拟AT(VAT)纳入注意力机制的概念,通过这种机制,即使从未标记的数据中也可以计算对抗性扰动。为了实现这种方法,我们提出了两种通用的训练技术,即注意力机制的增值税(Attention VAT)和注意力机制的“可解释”增值税(Attention iVAT),将注意力机制的AT扩展到半监督设置。特别是,Attention iVAT专注于注意力的差异;因此,它可以有效地学习更清晰的注意力并提高模型的可解释性,即使是未标记的数据。基于六个公共数据集的实证实验表明,我们的技术比传统的基于AT和基于VAT的技术提供了更好的预测性能,并且与人类在检测句子中重要单词时提供的证据具有更强的一致性。此外,我们的建议提供了这些优点,而不需要添加未标记数据的仔细选择。也就是说,即使使用我们基于VAT的技术的模型是在来自目标任务以外的源的未标记数据上训练的,预测性能和模型的可解释性都可以得到提高。
Although attention mechanisms have become fundamental components of deep learning models, they are vulnerable to perturbations, which may degrade the prediction performance and model interpretability. Adversarial training (AT) for attention mechanisms has successfully reduced such drawbacks by considering adversarial perturbations. However, this technique requires label information, and thus, its use is limited to supervised settings. In this study, we explore the concept of incorporating virtual AT (VAT) into the attention mechanisms, by which adversarial perturbations can be computed even from unlabeled data. To realize this approach, we propose two general training techniques, namely VAT for attention mechanisms (Attention VAT) and “interpretable” VAT for attention mechanisms (Attention iVAT), which extend AT for attention mechanisms to a semi-supervised setting. In particular, Attention iVAT focuses on the differences in attention; thus, it can efficiently learn clearer attention and improve model interpretability, even with unlabeled data. Empirical experiments based on six public datasets revealed that our techniques provide better prediction performance than conventional AT-based as well as VAT-based techniques, and stronger agreement with evidence that is provided by humans in detecting important words in sentences. Moreover, our proposal offers these advantages without needing to add the careful selection of unlabeled data. That is, even if the model using our VAT-based technique is trained on unlabeled data from a source other than the target task, both the prediction performance and model interpretability can be improved.