Efficient Adversarial Training With Transferable Adversarial Examples

Efficient Adversarial Training With Transferable Adversarial Examples
复制标题

DOI:
10.1109/cvpr42600.2020.00126
复制
发表时间:
2019-12
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Haizhong Zheng;Ziqi Zhang;Juncheng Gu;Honglak Lee;A. Prakash
Haizhong Zheng;Ziqi Zhang;Juncheng Gu;Honglak Lee;A. Prakash
中科院分区:
其他
文献类型:
--
作者:
Haizhong Zheng;Ziqi Zhang;Juncheng Gu;Honglak Lee;A. Prakash

文献摘要

被引文献

相似文献

对抗训练是一种保护分类模型免受对抗攻击的有效防御方法。但是,这种方法的一个局限性是,由于在训练过程中产生强大的对手实例的高成本,它可能需要额外的培训时间。在本文中,我们首先表明,在相同的训练过程中,模型之间的可传递性很高,即,一个时代的对抗性示例在随后的时期仍然是对抗性的。利用这一特性,我们提出了一种新颖的方法,具有可转移的对手实例(ATTA)的对抗训练,可以增强受过训练的模型的稳健性,并通过通过时期积累对抗性扰动来大大提高训练效率。与最先进的对抗训练方法相比,ATTA在CIFAR10上提高了对抗准确性高达7.2%,并且在MNIST和CIFAR10数据集上需要减少12〜14倍的训练时间,具有可比的模型鲁棒性。
Adversarial training is an effective defense method to protect classification models against adversarial attacks. However, one limitation of this approach is that it can require orders of magnitude additional training time due to high cost of generating strong adversarial examples during training. In this paper, we first show that there is high transferability between models from neighboring epochs in the same training process, i.e., adversarial examples from one epoch continue to be adversarial in subsequent epochs. Leveraging this property, we propose a novel method, Adversarial Training with Transferable Adversarial Examples (ATTA), that can enhance the robustness of trained models and greatly improve the training efficiency by accumulating adversarial perturbations through epochs. Compared to state-of-the-art adversarial training methods, ATTA enhances adversarial accuracy by up to 7.2% on CIFAR10 and requires 12~14x less training time on MNIST and CIFAR10 datasets with comparable model robustness.