Demystifying the Adversarial Robustness of Random Transformation Defenses

Demystifying the Adversarial Robustness of Random Transformation Defenses
复制标题

DOI:
10.48550/arxiv.2207.03574
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Chawin Sitawarin;Zachary Golan-Strieb;David A. Wagner
Chawin Sitawarin;Zachary Golan-Strieb;David A. Wagner
中科院分区:
其他
文献类型:
--
作者:
Chawin Sitawarin;Zachary Golan-Strieb;David A. Wagner

文献摘要

相似文献

神经网络对攻击缺乏鲁棒性,这引起了人们对自动驾驶汽车等安全敏感环境的担忧。虽然许多对策看起来很有希望,但只有少数能够经受住严格的评估。使用随机变换(RT)的防御已经显示出令人印象深刻的结果,特别是BaRT(Raff等人,2019年,在ImageNet。然而,这种类型的防御还没有得到严格的评估,使其鲁棒性知之甚少。它们的随机特性使得评估更具挑战性,并使许多针对确定性模型的攻击变得不适用。首先,我们证明了BPDA攻击(Athalye等人,2018a)是无效的,可能高估了其稳健性。然后,我们试图构建最强大的可能RT防御,通过知情的选择转换和贝叶斯优化调整其参数。此外,我们创建最强的攻击来评估我们的RT防御。我们的新攻击大大优于基线,与常用的EoT攻击的19%相比,准确率降低了83%(4)。3倍改善)。我们的研究结果表明,Imagenette数据集(ImageNet的十类子集)上的RT防御对对抗性示例并不鲁棒。进一步扩展研究,我们使用我们的新攻击来对抗训练RT防御(称为AdvRT),从而获得很大的鲁棒性增益。代码可在https://github.com/wagner-group/demystify-random-transform上获得。
Neural networks’ lack of robustness against attacks raises concerns in security-sensitive settings such as autonomous vehicles. While many coun-termeasures may look promising, only a few with-stand rigorous evaluation. Defenses using random transformations (RT) have shown impressive results, particularly BaRT (Raff et al., 2019) on ImageNet. However, this type of defense has not been rigorously evaluated, leaving its robustness properties poorly understood. Their stochastic properties make evaluation more challenging and render many proposed attacks on deterministic models inapplicable. First, we show that the BPDA attack (Athalye et al., 2018a) used in BaRT’s evaluation is ineffective and likely over-estimates its robustness. We then attempt to con-struct the strongest possible RT defense through the informed selection of transformations and Bayesian optimization for tuning their parameters. Furthermore, we create the strongest possible attack to evaluate our RT defense. Our new attack vastly outperforms the baseline, reducing the accuracy by 83% compared to the 19% re-duction by the commonly used EoT attack ( 4 . 3 × improvement). Our result indicates that the RT defense on Imagenette dataset (a ten-class subset of ImageNet) is not robust against adversarial examples. Extending the study further, we use our new attack to adversarially train RT defense (called AdvRT), resulting in a large robustness gain. Code is available at https://github.com/wagner-group/demystify-random-transform.