Gradient-based Analysis of NLP Models is Manipulable

Gradient-based Analysis of NLP Models is Manipulable
复制标题

DOI:
10.18653/v1/2020.findings-emnlp.24
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Junlin Wang;Jens Tuyls;Eric Wallace;Sameer Singh
Junlin Wang;Jens Tuyls;Eric Wallace;Sameer Singh
中科院分区:
其他
文献类型:
--
作者:
Junlin Wang;Jens Tuyls;Eric Wallace;Sameer Singh

文献摘要

相似文献

基于梯度的分析方法,如显著性图可视化和对抗性输入扰动,由于其简单、灵活,最重要的是,它们直接反映了模型内部,因此在解释神经NLP模型中得到了广泛的应用。然而,在本文中,我们证明了模型的梯度是容易操纵的,从而对基于梯度的分析的可靠性提出了质疑。特别地,我们将目标模型的层与Facade模型合并,Facade模型可以在不影响预测的情况下压倒梯度。这个Facade Model可以被训练成具有误导和与任务无关的梯度,例如只关注输入中的停止词。在各种NLP任务(情感分析、NLI和QA)上,我们表明合并模型有效地欺骗了不同的分析工具:显著性图与原始模型的显著性图不同,输入减少保留了更多不相关的输入令牌,对抗性扰动将不重要的令牌识别为非常重要的。
Gradient-based analysis methods, such as saliency map visualizations and adversarial input perturbations, have found widespread use in interpreting neural NLP models due to their simplicity, flexibility, and most importantly, the fact that they directly reflect the model internals. In this paper, however, we demonstrate that the gradients of a model are easily manipulable, and thus bring into question the reliability of gradient-based analyses. In particular, we merge the layers of a target model with a Facade Model that overwhelms the gradients without affecting the predictions. This Facade Model can be trained to have gradients that are misleading and irrelevant to the task, such as focusing only on the stop words in the input. On a variety of NLP tasks (sentiment analysis, NLI, and QA), we show that the merged model effectively fools different analysis tools: saliency maps differ significantly from the original model’s, input reduction keeps more irrelevant input tokens, and adversarial perturbations identify unimportant tokens as being highly important.