A General Framework for Adversarial Examples with Objectives

A General Framework for Adversarial Examples with Objectives
复制标题

DOI:
10.1145/3317611
复制
发表时间:
2019-07-01
影响因子:
2.3
通讯作者:
Reiter, Michael K.
Reiter, Michael K.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Sharif, Mahmood;Bhagavatula, Sruti;Reiter, Michael K.

文献摘要

被引文献

相似文献

被神经网络巧妙扰动而被错误分类的图像,称为对抗性示例,已经成为一个技术上的深层次挑战,也是几个应用领域的重要问题。大多数对抗性示例的研究都将扰动图像与原始图像相似作为其唯一的约束。然而,这些想法的实际应用通常需要示例满足额外的目标,这些目标通常通过对扰动过程的自定义修改来实现。在这篇文章中,我们提出了对抗生成网络(AGN),这是一种训练生成器神经网络的通用方法,可以生成满足预期目标的对抗示例。我们证明了活动星系核的能力,以适应广泛的目标,包括不精确的难以建模,在两个应用领域。特别是,我们展示了物理对抗性的例子-设计用于欺骗面部识别的框架-具有比以前的方法更好的鲁棒性,不显眼性和可扩展性,以及欺骗手写数字分类器的新攻击。
Images perturbed subtly to be misclassified by neural networks, called adversarial examples, have emerged as a technically deep challenge and an important concern for several application domains. Most research on adversarial examples takes as its only constraint that the perturbed images are similar to the originals. However, real-world application of these ideas often requires the examples to satisfy additional objectives, which are typically enforced through custom modifications of the perturbation process. In this article, we propose adversarial generative nets (AGNs), a general methodology to train a generator neural network to emit adversarial examples satisfying desired objectives. We demonstrate the ability of AGNs to accommodate a wide range of objectives, including imprecise ones difficult to model, in two application domains. In particular, we demonstrate physical adversarial examples-eyeglass frames designed to fool face recognition-with better robustness, inconspicuousness, and scalability than previous approaches, as well as a new attack to fool a handwritten-digit classifier.