NoiseCAM: Explainable AI for the Boundary Between Noise and Adversarial Attacks

NoiseCAM: Explainable AI for the Boundary Between Noise and Adversarial Attacks
复制标题

DOI:
10.1109/fuzz52849.2023.10309766
复制
发表时间:
2023-03
期刊:
2023 IEEE International Conference on Fuzzy Systems (FUZZ)
影响因子:
--
通讯作者:
Wen-Xi Tan;Justus Renkhoff;Alvaro Velasquez;Ziyu Wang;Lu Li;Jian Wang;Shuteng Niu;Fan Yang;Yongxin Liu;H. Song
Wen-Xi Tan;Justus Renkhoff;Alvaro Velasquez;Ziyu Wang;Lu Li;Jian Wang;Shuteng Niu;Fan Yang;Yongxin Liu;H. Song
中科院分区:
其他
文献类型:
--
作者:
Wen-Xi Tan;Justus Renkhoff;Alvaro Velasquez;Ziyu Wang;Lu Li;Jian Wang;Shuteng Niu;Fan Yang;Yongxin Liu;H. Song

文献摘要

相似文献

深度学习(DL)和深度神经网络(DNN)广泛用于各种领域。但是,对抗性攻击很容易误导神经网络并导致错误的决定。在安全应用中,防御机制高度优先。在本文中,首先,我们使用梯度类激活图(GradCAM)分析VGG-16网络的行为偏差,当时其输入与对抗性扰动或高斯噪声混合在一起。特别是,我们的方法可以找到对对抗扰动和高斯噪声敏感的脆弱层。我们还表明,脆弱层的行为偏差可用于检测对抗性示例。其次,我们提出了一种新型的NOISECAM算法,该算法整合了来自全球和像素水平加权类激活图的信息。我们的算法对对抗性扰动高度敏感,不会响应输入中混合的高斯随机噪声。第三,我们使用行为偏差和NOISECAM比较检测对抗示例,并且我们表明NOISECAM在其整体性能中的表现优于行为偏差模型。我们的工作可以提供有用的工具来防御对深神经网络的某些类型的对抗攻击。
Deep Learning (DL) and Deep Neural Networks (DNNs) are widely used in various domains. However, adversarial attacks can easily mislead a neural network and lead to wrong decisions. Defense mechanisms are highly preferred in safety- critical applications. In this paper, firstly, we use the gradient class activation map (GradCAM) to analyze the behavior deviation of the VGG-16 network when its inputs are mixed with adversarial perturbation or Gaussian noise. In particular, our method can locate vulnerable layers that are sensitive to adversarial perturbation and Gaussian noise. We also show that the behavior deviation of vulnerable layers can be used to detect adversarial examples. Secondly, we propose a novel NoiseCAM algorithm that integrates information from globally and pixel- level weighted class activation maps. Our algorithm is highly sensitive to adversarial perturbations and will not respond to Gaussian random noise mixed in the inputs. Third, we compare detecting adversarial examples using both behavior deviation and NoiseCAM, and we show that NoiseCAM outperforms behavior deviation modeling in its overall performance. Our work could provide a useful tool to defend against certain types of adversarial attacks on deep neural networks.