Universal Adversarial Training

Universal Adversarial Training
复制标题

DOI:
10.1609/aaai.v34i04.6017
复制
发表时间:
2018-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Ali Shafahi;Mahyar Najibi;Zheng Xu;John P. Dickerson;L. Davis;T. Goldstein
Ali Shafahi;Mahyar Najibi;Zheng Xu;John P. Dickerson;L. Davis;T. Goldstein
中科院分区:
其他
文献类型:
--
作者:
Ali Shafahi;Mahyar Najibi;Zheng Xu;John P. Dickerson;L. Davis;T. Goldstein

文献摘要

被引文献

相似文献

标准对抗攻击通过向像素添加专门定制的小扰动来改变所选图像的预测类别标签。相比之下,通用扰动是可以添加到广泛类别的图像中的任何图像的更新,同时仍然改变预测的类别标签。我们研究了通用对抗性扰动的有效生成,以及增强网络抵御这些攻击的有效方法。我们提出了一种简单的基于优化的通用攻击,它可以将ImageNet上各种网络架构的top-1准确率降低到20%以下,同时学习通用扰动的速度比标准方法快13倍。为了抵御这些扰动,我们提出了通用对抗训练,它将鲁棒分类器生成问题建模为两个玩家的最小-最大游戏,并且生成鲁棒的模型,其成本仅为自然训练的2倍。我们还提出了一种几乎不需要额外计算的同步随机梯度方法,这使得我们可以在ImageNet上进行通用的对抗训练。
Standard adversarial attacks change the predicted class label of a selected image by adding specially tailored small perturbations to its pixels. In contrast, a universal perturbation is an update that can be added to any image in a broad class of images, while still changing the predicted class label. We study the efficient generation of universal adversarial perturbations, and also efficient methods for hardening networks to these attacks. We propose a simple optimization-based universal attack that reduces the top-1 accuracy of various network architectures on ImageNet to less than 20%, while learning the universal perturbation 13× faster than the standard method.To defend against these perturbations, we propose universal adversarial training, which models the problem of robust classifier generation as a two-player min-max game, and produces robust models with only 2× the cost of natural training. We also propose a simultaneous stochastic gradient method that is almost free of extra computation, which allows us to do universal adversarial training on ImageNet.