Ensemble transfer attack targeting text classification systems

Ensemble transfer attack targeting text classification systems
复制标题

针对文本分类系统的集成转移攻击

DOI:
10.1016/j.cose.2022.102695
复制
发表时间:
2022
期刊:
Comput. Secur.
影响因子:
--
通讯作者:
["Hyung
["Hyung
中科院分区:
--
文献类型:
--
作者:
["Hyung

文献摘要

被引文献

相似文献

深度神经网络为图像识别、语音识别、文本识别和模式识别提供了良好的性能。然而,这样的网络很容易受到对抗性例子的攻击。对抗性样本是通过向原始样本添加少量噪声来创建的,以这种方式,人类不会察觉到任何问题,但样本将被分类模型错误地分类。对抗性示例主要在图像背景下进行研究,但研究已扩展到包括文本领域。在文本上下文中,对抗性示例是文本样本,其中某些重要的单词已经被更改,使得样本将被模型错误分类,即使对人类来说,它在含义和语法方面与原始文本相同。然而,使用文本对抗示例进行黑盒攻击的研究很少。在本文中,我们提出了系综转移textfooler方法。该方法在生成一个同时欺骗多个模型的集成对抗性示例后,对未知模型进行黑盒攻击。实验使用电影评论数据集进行,并使用TensorFlow作为机器学习库。实验结果表明,该方法的攻击成功率为71.64%,而传统的传输攻击的攻击成功率分别为19.01%,24.29%和44.96%,使用对抗性示例来欺骗WordCNN,WordLSTM和BERT模型。
Deep neural networks provide good performance for image recognition, speech recognition, text recognition, and pattern recognition. However, such networks are vulnerable to attack by adversarial examples. Adversarial examples are created by adding a small amount of noise to an original sample in such a way that no problem is perceptible to humans yet the sample will be incorrectly classified by a classification model. Adversarial examples have been studied mainly in the context of images, but research has expanded to include the text domain. In the textual context, an adversarial example is a sample of text in which certain important words have been changed so that the sample will be misclassified by a model even though to humans it is the same as the original text in terms of meaning and grammar. However, studies of black box attacks using text adversarial examples are sparse. In this paper, we propose the ensemble transfer textfooler method. This method performs a black box attack on an unknown model after generating an ensemble adversarial example that simultaneously deceives several models. Experiments were conducted using a movie review dataset and with TensorFlow as the machine learning library. The experimental results show that the proposed method has an attack success rate of 71.64%, in contrast to the 19.01%, 24.29%, and 44.96% attack success rate for the conventional transfer attacks using an adversarial example generated to deceive a WordCNN, WordLSTM, and BERT model.