Adversarial Training and Robustness for Multiple Perturbations

Adversarial Training and Robustness for Multiple Perturbations
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
--
影响因子:
--
通讯作者:
Florian Tramèr;D. Boneh
Florian Tramèr;D. Boneh
中科院分区:
其他
文献类型:
--
作者:
Florian Tramèr;D. Boneh

文献摘要

被引文献

相似文献

对抗性示例的防御,例如对抗性训练,通常针对单个扰动类型(例如,small $\ell_\infty$-noise)。对于其他扰动,这些防御措施不能提供任何保证,有时甚至会增加模型的脆弱性。我们的目标是了解这种鲁棒性权衡的原因,并训练同时对多种扰动类型具有鲁棒性的模型。我们证明了权衡不同类型的$\ell_p$有界和空间扰动的鲁棒性必须存在于一个自然和简单的统计设置。我们证实了我们的正式分析,证明类似的鲁棒性权衡MNIST和CIFAR 10。建立在新的多扰动对抗训练方案,和一种新的有效的攻击,寻找$\ell_1$-有界对抗的例子,我们表明,没有模型训练对多个攻击实现鲁棒性的竞争与模型训练的每个攻击单独。特别是,我们发现了MNIST上一个有害的梯度掩蔽现象,它导致对抗训练与一阶$\ell_\infty,\ell_1$和$\ell_2$对手,以达到只有$50\%$的准确性。我们的研究结果质疑扩展对抗鲁棒性和对抗训练到多种扰动类型的可行性和计算可扩展性。
Defenses against adversarial examples, such as adversarial training, are typically tailored to a single perturbation type (e.g., small $\ell_\infty$-noise). For other perturbations, these defenses offer no guarantees and, at times, even increase the model's vulnerability. Our aim is to understand the reasons underlying this robustness trade-off, and to train models that are simultaneously robust to multiple perturbation types. We prove that a trade-off in robustness to different types of $\ell_p$-bounded and spatial perturbations must exist in a natural and simple statistical setting. We corroborate our formal analysis by demonstrating similar robustness trade-offs on MNIST and CIFAR10. Building upon new multi-perturbation adversarial training schemes, and a novel efficient attack for finding $\ell_1$-bounded adversarial examples, we show that no model trained against multiple attacks achieves robustness competitive with that of models trained on each attack individually. In particular, we uncover a pernicious gradient-masking phenomenon on MNIST, which causes adversarial training with first-order $\ell_\infty, \ell_1$ and $\ell_2$ adversaries to achieve merely $50\%$ accuracy. Our results question the viability and computational scalability of extending adversarial robustness, and adversarial training, to multiple perturbation types.