NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks

NATTACK: Learning the Distributions of Adversarial Examples for an Improved Black-Box Attack on Deep Neural Networks
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yandong Li;Lijun Li;Liqiang Wang;Tong Zhang;Boqing Gong
Yandong Li;Lijun Li;Liqiang Wang;Tong Zhang;Boqing Gong
中科院分区:
其他
文献类型:
--
作者:
Yandong Li;Lijun Li;Liqiang Wang;Tong Zhang;Boqing Gong

文献摘要

被引文献

相似文献

强大的对抗性攻击方法对于理解如何构建强大的深度神经网络(DNN)和彻底测试防御技术至关重要。在本文中,我们提出了一种黑盒对抗性攻击算法,该算法可以击败普通 DNN 和最近开发的各种防御技术生成的算法。我们的算法不是为目标 DNN 的良性输入搜索“最佳”对抗性示例,而是找到以输入为中心的小区域上的概率密度分布,这样从该分布中抽取的样本可能是对抗性示例,而无需访问 DNN 的内部层或权重。我们的方法是通用的,因为它可以通过单一算法成功攻击不同的神经网络。它也很强大;根据对 2 个普通 DNN 和 13 个防御 DNN 的测试,它在大多数测试用例中都优于最先进的黑盒或白盒攻击方法。此外,我们的结果表明,对抗性训练仍然是最好的防御技术之一,并且对抗性示例在受防御的 DNN 之间的转移不如在普通 DNN 之间转移。
Powerful adversarial attack methods are vital for understanding how to construct robust deep neural networks (DNNs) and for thoroughly testing defense techniques. In this paper, we propose a black-box adversarial attack algorithm that can defeat both vanilla DNNs and those generated by various defense techniques developed recently. Instead of searching for an "optimal" adversarial example for a benign input to a targeted DNN, our algorithm finds a probability density distribution over a small region centered around the input, such that a sample drawn from this distribution is likely an adversarial example, without the need of accessing the DNN's internal layers or weights. Our approach is universal as it can successfully attack different neural networks by a single algorithm. It is also strong; according to the testing against 2 vanilla DNNs and 13 defended ones, it outperforms state-of-the-art black-box or white-box attack methods for most test cases. Additionally, our results reveal that adversarial training remains one of the best defense techniques, and the adversarial examples are not as transferable across defended DNNs as them across vanilla DNNs.