Semantic Adversarial Attacks: Parametric Transformations That Fool Deep Classifiers

Semantic Adversarial Attacks: Parametric Transformations That Fool Deep Classifiers
复制标题

DOI:
10.1109/iccv.2019.00487
复制
发表时间:
2019-04
期刊:
2019 IEEE/CVF International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
Ameya Joshi;Amitangshu Mukherjee;S. Sarkar;C. Hegde
Ameya Joshi;Amitangshu Mukherjee;S. Sarkar;C. Hegde
中科院分区:
其他
文献类型:
--
作者:
Ameya Joshi;Amitangshu Mukherjee;S. Sarkar;C. Hegde

文献摘要

相似文献

深度神经网络已被证明对被不可感知的扰动破坏的对抗性输入图像表现出有趣的脆弱性。然而,大多数对抗性攻击都假设对图像像素空间进行全局的细粒度控制。在本文中,我们考虑一个不同的设置:如果对手只能改变输入图像的特定属性会发生什么?这些将生成可能明显不同的输入,但仍然看起来很自然,足以欺骗分类器。我们提出了一种新的方法来生成这样的“语义”对抗的例子,通过优化一个特定的对抗损失的范围空间的参数条件生成模型。我们展示了对在人脸图像上训练的二进制分类器的攻击的实现,并表明存在这种看起来自然的语义对抗示例。我们评估我们的攻击的有效性合成和真实的数据,并提出了详细的比较与现有的攻击方法。我们补充我们的经验结果与理论界证明存在这样的参数对抗的例子。
Deep neural networks have been shown to exhibit an intriguing vulnerability to adversarial input images corrupted with imperceptible perturbations. However, the majority of adversarial attacks assume global, fine-grained control over the image pixel space. In this paper, we consider a different setting: what happens if the adversary could only alter specific attributes of the input image? These would generate inputs that might be perceptibly different, but still natural-looking and enough to fool a classifier. We propose a novel approach to generate such ``semantic'' adversarial examples by optimizing a particular adversarial loss over the range-space of a parametric conditional generative model. We demonstrate implementations of our attacks on binary classifiers trained on face images, and show that such natural-looking semantic adversarial examples exist. We evaluate the effectiveness of our attack on synthetic and real data, and present detailed comparisons with existing attack methods. We supplement our empirical results with theoretical bounds that demonstrate the existence of such parametric adversarial examples.