On the Transferability of Adversarial Attacks against Neural Text Classifier

On the Transferability of Adversarial Attacks against Neural Text Classifier
复制标题

DOI:
10.18653/v1/2021.emnlp-main.121
复制
发表时间:
2020-11
期刊:
--
影响因子:
--
通讯作者:
Liping Yuan;Xiaoqing Zheng;Yi Zhou;Cho-Jui Hsieh;Kai-Wei Chang
Liping Yuan;Xiaoqing Zheng;Yi Zhou;Cho-Jui Hsieh;Kai-Wei Chang
中科院分区:
其他
文献类型:
--
作者:
Liping Yuan;Xiaoqing Zheng;Yi Zhou;Cho-Jui Hsieh;Kai-Wei Chang

文献摘要

相似文献

深度神经网络很容易受到对抗性攻击,输入的微小扰动就会改变模型的预测。在许多情况下,为一个模型故意制作的恶意输入可能会欺骗另一个模型。在本文中,我们提出了第一项系统研究文本分类模型的对抗性示例的可迁移性的研究,并探讨了包括网络架构、标记化方案、词嵌入和模型容量在内的各种因素如何影响对抗性示例的可迁移性。基于这些研究,我们提出了一种遗传算法来寻找模型集合,该模型集合可用于诱导对抗性示例来欺骗几乎所有现有模型。这样的对抗性例子反映了学习过程的缺陷和训练集中的数据偏差。最后,我们从这些对抗性示例中得出可用于模型诊断的单词替换规则。
Deep neural networks are vulnerable to adversarial attacks, where a small perturbation to an input alters the model prediction. In many cases, malicious inputs intentionally crafted for one model can fool another model. In this paper, we present the first study to systematically investigate the transferability of adversarial examples for text classification models and explore how various factors, including network architecture, tokenization scheme, word embedding, and model capacity, affect the transferability of adversarial examples. Based on these studies, we propose a genetic algorithm to find an ensemble of models that can be used to induce adversarial examples to fool almost all existing models. Such adversarial examples reflect the defects of the learning process and the data bias in the training set. Finally, we derive word replacement rules that can be used for model diagnostics from these adversarial examples.