Minimally Distorted Structured Adversarial Attacks

Minimally Distorted Structured Adversarial Attacks
复制标题

DOI:
10.1007/s11263-022-01701-w
复制
发表时间:
2022-10
影响因子:
19.5
通讯作者:
Ehsan Kazemi;Thomas Kerdreux;Liqiang Wang
Ehsan Kazemi;Thomas Kerdreux;Liqiang Wang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ehsan Kazemi;Thomas Kerdreux;Liqiang Wang

文献摘要

相似文献

白色盒对抗扰动是通过迭代优化算法产生的,最常见的是通过最小化原始图像邻域上的对抗损失,即所谓的失真集。用不同的规范约束对抗性搜索会产生结构合理的对抗性示例。在这里,我们探索几个失真集与结构增强算法。对抗性示例的这些新结构可能会为可证明和经验的鲁棒机制带来挑战。由于对抗鲁棒性仍然是一个经验领域,防御机制也应该合理地评估不同结构的攻击。此外,这些结构化对抗扰动可以允许比它们的对应物更大的失真大小,同时保持不可感知或可感知的图像的自然失真。我们将在这项工作中证明,所提出的结构化对抗性示例可以显着降低对抗性训练分类器的分类精度,同时显示出低失真率。例如,在ImagNet数据集上,结构化攻击将对抗模型的准确性降低到接近零,只有50%的失真是使用PGD等白盒攻击产生的。作为一个副产品,我们在结构化对抗性示例上的发现可以用于模型的对抗性正则化,以使模型更鲁棒或提高其在结构不同的数据集上的泛化性能。
White box adversarial perturbations are generated via iterative optimization algorithms most often by minimizing an adversarial loss on aneighborhood of the original image, the so-called distortion set. Constraining the adversarial search with different norms results in disparately structured adversarial examples. Here we explore several distortion sets with structure-enhancing algorithms. These new structures for adversarial examples might provide challenges for provable and empirical robust mechanisms. Because adversarial robustness is still an empirical field, defense mechanisms should also reasonably be evaluated against differently structured attacks. Besides, these structured adversarial perturbations may allow for larger distortions size than theircounterpart while remaining imperceptible or perceptible as natural distortions of the image. We will demonstrate in this work that the proposed structured adversarial examples can significantly bring down the classification accuracy of adversarially trained classifiers while showing a lowdistortion rate. For instance, on ImagNet dataset the structured attacks drop the accuracy of the adversarial model to near zero with only 50% ofdistortion generated using white-box attacks like PGD. As a byproduct, our findings on structured adversarial examples can be used for adversarial regularization of models to make models more robust or improve their generalization performance on datasets that are structurally different.