Explicit Tradeoffs between Adversarial and Natural Distributional Robustness

Explicit Tradeoffs between Adversarial and Natural Distributional Robustness
复制标题

DOI:
10.48550/arxiv.2209.07592
复制
发表时间:
2022-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Mazda Moayeri;Kiarash Banihashem;S. Feizi
Mazda Moayeri;Kiarash Banihashem;S. Feizi
中科院分区:
其他
文献类型:
--
作者:
Mazda Moayeri;Kiarash Banihashem;S. Feizi

文献摘要

被引文献

相似文献

已有的几个工作分别研究了深层神经网络的对抗性或自然分布稳健性。然而,在实践中,模型需要同时具备这两种类型的健壮性,以确保可靠性。在这项工作中,我们弥合了这一差距,并表明事实上,在对抗性分布健壮性和自然分布健壮性之间存在明显的权衡。我们首先考虑具有不相交的核心和伪特征集的高斯数据的简单线性回归设置。在这种情况下,通过理论和实证分析,我们发现:(1)对抗性训练增加了对虚假特征的依赖;(2)对于对抗性训练,只有当虚假特征的规模大于核心特征的规模时,才会产生虚假依赖;(3)对抗性训练可能会产生意想不到的结果,降低分布的稳健性,特别是当虚假相关性在新的测试域中发生变化时。接下来,我们使用在五个基准数据集(ObjectNet,RIVAL10,Salient ImageNet-1M,ImageNet-9,Water Birds)上评估的20个对抗性训练模型的测试套件,提供了广泛的经验证据,表明对抗性训练的分类器比标准训练的分类器更依赖背景,验证了我们的理论结果。我们还证明了训练数据中的虚假相关性(当保留在测试域中)可以提高对手的稳健性,揭示了以前关于对手脆弱性植根于虚假相关性的说法是不完整的。
Several existing works study either adversarial or natural distributional robustness of deep neural networks separately. In practice, however, models need to enjoy both types of robustness to ensure reliability. In this work, we bridge this gap and show that in fact, explicit tradeoffs exist between adversarial and natural distributional robustness. We first consider a simple linear regression setting on Gaussian data with disjoint sets of core and spurious features. In this setting, through theoretical and empirical analysis, we show that (i) adversarial training with $\ell_1$ and $\ell_2$ norms increases the model reliance on spurious features; (ii) For $\ell_\infty$ adversarial training, spurious reliance only occurs when the scale of the spurious features is larger than that of the core features; (iii) adversarial training can have an unintended consequence in reducing distributional robustness, specifically when spurious correlations are changed in the new test domain. Next, we present extensive empirical evidence, using a test suite of twenty adversarially trained models evaluated on five benchmark datasets (ObjectNet, RIVAL10, Salient ImageNet-1M, ImageNet-9, Waterbirds), that adversarially trained classifiers rely on backgrounds more than their standardly trained counterparts, validating our theoretical results. We also show that spurious correlations in training data (when preserved in the test domain) can improve adversarial robustness, revealing that previous claims that adversarial vulnerability is rooted in spurious correlations are incomplete.