Priority Adversarial Example in Evasion Attack on Multiple Deep Neural Networks

Priority Adversarial Example in Evasion Attack on Multiple Deep Neural Networks
复制标题

多个深度神经网络规避攻击中的优先对抗示例

DOI:
10.1109/icaiic.2019.8669034
复制
发表时间:
2019
期刊:
2019 International Conference on Artificial Intelligence in Information and Communication (ICAIIC)
影响因子:
--
通讯作者:
D. Choi
D. Choi
中科院分区:
--
文献类型:
--
作者:
Hyun Kwon;H. Yoon;D. Choi

文献摘要

被引文献

相似文献

深度神经网络(DNN)为机器学习任务(如图像识别、语音识别、模式识别和入侵检测)提供了上级性能。然而,通过向原始数据添加一点噪声而创建的对抗性示例可能会导致DNN的错误分类,并且人眼无法检测到与原始数据的差异。例如,如果攻击者生成被DNN错误分类的修改的左转道路标志,则具有DNN的自动驾驶车辆将不正确地将修改的左转道路标志分类为右转标志,而人类将正确地将修改的标志分类为左转标志。这样一个对抗性的例子对DNN来说是一个严重的威胁。最近,引入了一个多目标对抗性示例,该示例使用单个修改后的图像在每个目标类别内导致多个模型的错误分类。但是,它存在的漏洞是,随着目标模型数量的增加,整体攻击成功率会降低。因此,如果存在攻击者希望针对的多个模型,则攻击者需要通过考虑每个模型的攻击优先级来控制每个模型的攻击成功率。在本文中,我们提出了一个优先级对抗的例子,它考虑了针对多个模型的情况下每个模型的攻击优先级。该方法通过在生成过程中调整攻击函数的权重来控制每个模型的攻击成功率,同时保持最小失真。我们使用Tensorflow,一个广泛使用的机器学习库,和MNIST作为数据集。实验结果表明,该方法可以通过考虑每个模型的攻击优先级来控制每个模型的攻击成功率,同时保持最小的失真(在有针对性和无针对性的攻击中,平均分别为3.95和2.45)。
Deep neural networks (DNNs) provide superior per-formance on machine learning tasks such as image recognition, speech recognition, pattern recognition, and intrusion detection. However, an adversarial example created by adding a little noise to the original data can lead to misclassification by the DNN, and the human eye cannot detect the difference from the original data. For example, if an attacker generates a modified left-turn road sign to be incorrectly categorized by a DNN, an autonomous vehicle with the DNN will incorrect classify the modified left-turn road sign as a right-turn sign, whereas a human will correctly classify the modified sign as a left-turn sign. Such an adversarial example is a serious threat to a DNN. Recently, a multi-target adversarial example was introduced that causes misclassification by several models within each target class using a single modified image. However, it has the vulnerability that as the number of target models increases, the overall attack success rate is reduced. Therefore, if there are several models that the attacker wishes to target, the attacker needs to control the attack success rate for each model by considering the attack priority for each model. In this paper, we propose a priority adversarial example that considers the attack priority for each model in cases targeting several models. The proposed method controls the attack success rate for each model by adjusting the weight of the attack function in the generation process, while maintaining minimum distortion. We used Tensorflow, a widely used machine learning library, and MNIST as the dataset. Experimental results show that the proposed method can control the attack success rate for each model by considering the attack priority of each model while maintaining minimum distortion (on average 3.95 and 2.45 in targeted and untargeted attacks, respectively).