Towards Evaluating the Robustness of Neural Networks Learned by Transduction

Towards Evaluating the Robustness of Neural Networks Learned by Transduction
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jiefeng Chen;Xi Wu;Yang Guo;Yingyu Liang;S. Jha
Jiefeng Chen;Xi Wu;Yang Guo;Yingyu Liang;S. Jha
中科院分区:
其他
文献类型:
--
作者:
Jiefeng Chen;Xi Wu;Yang Guo;Yingyu Liang;S. Jha

文献摘要

被引文献

相似文献

人们对将转导学习用于对抗鲁棒性的兴趣日益浓厚(Goldwasser等人,NeurIPS 2020;Wu等人,ICML 2020;Wang等人,ArXiv 2021)。与传统防御方法相比,这些防御机制基于测试时的输入“动态学习”模型;从理论上讲,攻击这些防御可归结为解决一个双层优化问题,这给构造自适应攻击带来了困难。在本文中,我们从有原则的威胁分析角度研究这些防御机制。我们为基于转导学习的防御制定并分析了威胁模型,并指出了一些重要的细微之处。我们提出了攻击模型空间以解决双层攻击目标的原则,并提出了贪婪模型空间攻击(GMSA),这是一个可作为评估基于转导学习的防御的新基准的攻击框架。通过系统评估,我们表明,即使是弱实例化的GMSA也能够突破之前基于转导学习的防御,这些防御对之前的攻击(如AutoAttack)具有抵抗力。从积极的方面来看,我们报告了一个有点令人惊讶的“转导对抗训练”的实证结果:在测试时使用新的随机性对模型进行对抗性再训练,会显著提高针对我们所考虑的攻击的鲁棒性。
There has been emerging interest in using transductive learning for adversarial robustness (Goldwasser et al., NeurIPS 2020; Wu et al., ICML 2020; Wang et al., ArXiv 2021). Compared to traditional defenses, these defense mechanisms"dynamically learn"the model based on test-time input; and theoretically, attacking these defenses reduces to solving a bilevel optimization problem, which poses difficulty in crafting adaptive attacks. In this paper, we examine these defense mechanisms from a principled threat analysis perspective. We formulate and analyze threat models for transductive-learning based defenses, and point out important subtleties. We propose the principle of attacking model space for solving bilevel attack objectives, and present Greedy Model Space Attack (GMSA), an attack framework that can serve as a new baseline for evaluating transductive-learning based defenses. Through systematic evaluation, we show that GMSA, even with weak instantiations, can break previous transductive-learning based defenses, which were resilient to previous attacks, such as AutoAttack. On the positive side, we report a somewhat surprising empirical result of"transductive adversarial training": Adversarially retraining the model using fresh randomness at the test time gives a significant increase in robustness against attacks we consider.