Exploring Adversarial Attacks on Neural Networks: An Explainable Approach

Exploring Adversarial Attacks on Neural Networks: An Explainable Approach
复制标题

DOI:
10.1109/ipccc55026.2022.9894322
复制
发表时间:
2022-11
期刊:
2022 IEEE International Performance, Computing, and Communications Conference (IPCCC)
影响因子:
--
通讯作者:
Justus Renkhoff;Wenkai Tan;Alvaro Velasquez;William Yichen Wang;Yongxin Liu;Jian Wang;Shuteng Niu;L. Fazlic;Guido Dartmann;H. Song
Justus Renkhoff;Wenkai Tan;Alvaro Velasquez;William Yichen Wang;Yongxin Liu;Jian Wang;Shuteng Niu;L. Fazlic;Guido Dartmann;H. Song
中科院分区:
其他
文献类型:
--
作者:
Justus Renkhoff;Wenkai Tan;Alvaro Velasquez;William Yichen Wang;Yongxin Liu;Jian Wang;Shuteng Niu;L. Fazlic;Guido Dartmann;H. Song

文献摘要

相似文献

深度学习(DL)正在应用于各个领域,特别是在自动驾驶等安全关键应用中。因此,确保这些方法的鲁棒性,从而抵抗对抗性攻击引起的不确定行为具有重要意义。本文使用梯度热图分析了VGG-16模型在输入图像中混合对抗性噪声和统计相似的高斯随机噪声时的响应特性。特别是,我们逐层比较网络响应,以确定错误发生的位置。几个有趣的发现派生。首先,与高斯随机噪声相比,故意生成的对抗性噪声通过分散网络中的集中区域而导致严重的行为偏差。其次,在许多情况下,对抗性示例只需要妥协几个中间块就可以误导最终决策。第三,我们的实验表明,特定的块更容易受到攻击,更容易被对抗性的例子利用。最后,我们证明了VGG-16模型的Block4_conv1和Block 5_cov 1层更容易受到对抗性攻击。我们的工作可能为开发更可靠的深度神经网络(DNN)模型提供有用的见解。
Deep Learning (DL) is being applied in various domains, especially in safety-critical applications such as autonomous driving. Consequently, it is of great significance to ensure the robustness of these methods and thus counteract uncertain behaviors caused by adversarial attacks. In this paper, we use gradient heatmaps to analyze the response characteristics of the VGG-16 model when the input images are mixed with adversarial noise and statistically similar Gaussian random noise. In particular, we compare the network response layer by layer to determine where errors occurred. Several interesting findings are derived. First, compared to Gaussian random noise, intentionally generated adversarial noise causes severe behavior deviation by distracting the area of concentration in the networks. Second, in many cases, adversarial examples only need to compromise a few intermediate blocks to mislead the final decision. Third, our experiments revealed that specific blocks are more vulnerable and easier to exploit by adversarial examples. Finally, we demonstrate that the layers Block4_conv1 and Block5_ cov1 of the VGG-16 model are more susceptible to adversarial attacks. Our work could potentially provide useful insights into developing more reliable Deep Neural Network (DNN) models.