Generation of Low Distortion Adversarial Attacks via Convex Programming

Generation of Low Distortion Adversarial Attacks via Convex Programming
复制标题

DOI:
10.1109/icdm.2019.00195
复制
发表时间:
2019-11
期刊:
2019 IEEE International Conference on Data Mining (ICDM)
影响因子:
--
通讯作者:
Tianyun Zhang;Sijia Liu;Yanzhi Wang;M. Fardad
Tianyun Zhang;Sijia Liu;Yanzhi Wang;M. Fardad
中科院分区:
其他
文献类型:
--
作者:
Tianyun Zhang;Sijia Liu;Yanzhi Wang;M. Fardad

文献摘要

相似文献

随着深度神经网络(DNNS)在各种任务中实现非凡的表现,在对抗攻击下测试其稳健性变得至关重要。对抗性攻击,也称为对抗性示例,用于测量DNN的鲁棒性,并通过将不可察觉的扰动纳入输入数据来生成,以改变DNN的分类。在该领域的先前工作中,大多数基于优化的方法采用梯度下降来找到对抗性示例。在本文中,我们提出了一种创新方法,该方法通过凸编程生成对抗性示例。我们的实验结果表明,与C&W攻击相比,我们可以生成具有较低失真且可传递性较高的对抗性示例,C&W攻击是DNN的当前最新对抗攻击方法。我们在原始的未防御模型和对抗训练的模型上都达到了100%的攻击成功率。对于最佳情况和CIFAR-10数据集的最佳情况和平均情况,我们对L_INF攻击的扭曲分别比C&W攻击低31%和18%。
As deep neural networks (DNNs) achieve extraordinary performance in a wide range of tasks, testing their robustness under adversarial attacks becomes paramount. Adversarial attacks, also known as adversarial examples, are used to measure the robustness of DNNs and are generated by incorporating imperceptible perturbations into the input data with the intention of altering a DNN's classification. In prior work in this area, most of the proposed optimization based methods employ gradient descent to find adversarial examples. In this paper, we present an innovative method which generates adversarial examples via convex programming. Our experiment results demonstrate that we can generate adversarial examples with lower distortion and higher transferability than the C&W attack, which is the current state-of-the-art adversarial attack method for DNNs. We achieve 100% attack success rate on both the original undefended models and the adversarially-trained models. Our distortions of the L_inf attack are respectively 31% and 18% lower than the C&W attack for the best case and average case on the CIFAR-10 data set.