Closer Look at the Transferability of Adversarial Examples: How They Fool Different Models Differently

Closer Look at the Transferability of Adversarial Examples: How They Fool Different Models Differently
复制标题

DOI:
10.1109/wacv56688.2023.00141
复制
发表时间:
2021-12
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Futa Waseda;Sosuke Nishikawa;Trung-Nghia Le;H. Nguyen;I. Echizen
Futa Waseda;Sosuke Nishikawa;Trung-Nghia Le;H. Nguyen;I. Echizen
中科院分区:
其他
文献类型:
--
作者:
Futa Waseda;Sosuke Nishikawa;Trung-Nghia Le;H. Nguyen;I. Echizen

文献摘要

相似文献

深度神经网络容易受到对抗性示例(AE)的影响,这些示例具有对抗性可转移性:为源模型生成的AE可能会误导另一个(目标)模型的预测。然而,在哪类目标模型的预测被误导的方面,可转移性还没有被理解(即,类感知可转移性)。在本文中,我们区分的情况下,目标模型预测相同的错误类作为源模型(“相同的错误”)或不同的错误类(“不同的错误”)进行分析,并提供一个机制的解释。我们发现:(1)AE倾向于导致相同的错误,这与“非目标可转移性”相关;然而,(2)即使在相似的模型之间,也会出现不同的错误,无论扰动大小如何。此外,我们提出的证据表明,相同的错误和不同的错误之间的差异可以解释为非鲁棒性的功能,预测,但人类无法解释的模式:不同的错误发生时,AE中的非鲁棒性功能不同的模型使用。因此,非鲁棒特征可以为AE的类感知可转移性提供一致的解释。
Deep neural networks are vulnerable to adversarial examples (AEs), which have adversarial transferability: AEs generated for the source model can mislead another (target) model’s predictions. However, the transferability has not been understood in terms of to which class target model’s predictions were misled (i.e., class-aware transferability). In this paper, we differentiate the cases in which a target model predicts the same wrong class as the source model ("same mistake") or a different wrong class ("different mistake") to analyze and provide an explanation of the mechanism. We find that (1) AEs tend to cause same mistakes, which correlates with "non-targeted transferability"; how-ever, (2) different mistakes occur even between similar models, regardless of the perturbation size. Furthermore, we present evidence that the difference between same mistakes and different mistakes can be explained by non-robust features, predictive but human-uninterpretable patterns: different mistakes occur when non-robust features in AEs are used differently by models. Non-robust features can thus provide consistent explanations for the class-aware transferability of AEs.