ADMM attack: an enhanced adversarial attack for deep neural networks with undetectable distortions

ADMM attack: an enhanced adversarial attack for deep neural networks with undetectable distortions
复制标题

DOI:
10.1145/3287624.3288750
复制
发表时间:
2019-01
期刊:
Proceedings of the 24th Asia and South Pacific Design Automation Conference
影响因子:
--
通讯作者:
Pu Zhao;Kaidi Xu;Sijia Liu;Yanzhi Wang;X. Lin
Pu Zhao;Kaidi Xu;Sijia Liu;Yanzhi Wang;X. Lin
中科院分区:
其他
文献类型:
--
作者:
Pu Zhao;Kaidi Xu;Sijia Liu;Yanzhi Wang;X. Lin

文献摘要

相似文献

最近的许多研究表明,最先进的深度神经网络(DNN)可能很容易被对抗性的例子所欺骗,这些例子是通过对抗性攻击在原始法律的输入上添加精心制作的和视觉上难以察觉的扭曲而生成的。对抗性示例可能会导致DNN将其错误分类为任何目标标签。在文献中,提出了各种方法来最小化失真的不同lp范数。然而,缺乏一个适用于所有类型对抗性攻击的通用框架。为了更好地理解DNN的安全属性,我们提出了一个构建对抗性示例的通用框架,该框架利用交替方向乘法(ADMM)来分割优化方法,以有效地最小化失真的各种lp范数,包括l0,l1,l2和l∞范数。因此,所提出的一般框架统一了制作l0,l1,l2和l∞攻击的方法。实验结果表明,与现有的攻击方法相比,本文提出的ADMM攻击具有较高的攻击成功率和最小的误分类失真.
Many recent studies demonstrate that state-of-the-art Deep neural networks (DNNs) might be easily fooled by adversarial examples, generated by adding carefully crafted and visually imperceptible distortions onto original legal inputs through adversarial attacks. Adversarial examples can lead the DNN to misclassify them as any target labels. In the literature, various methods are proposed to minimize the different lp norms of the distortion. However, there lacks a versatile framework for all types of adversarial attacks. To achieve a better understanding for the security properties of DNNs, we propose a general framework for constructing adversarial examples by leveraging Alternating Direction Method of Multipliers (ADMM) to split the optimization approach for effective minimization of various lp norms of the distortion, including l0, l1, l2, and l∞ norms. Thus, the proposed general framework unifies the methods of crafting l0, l1, l2, and l∞ attacks. The experimental results demonstrate that the proposed ADMM attacks achieve both the high attack success rate and the minimal distortion for the misclassification compared with state-of-the-art attack methods.