Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations

Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial Perturbations
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
Proceedings of the 2019 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Florian Tramèr;Jens Behrmann;Nicholas Carlini;Nicolas Papernot;J. Jacobsen
Florian Tramèr;Jens Behrmann;Nicholas Carlini;Nicolas Papernot;J. Jacobsen
中科院分区:
其他
文献类型:
--
作者:
Florian Tramèr;Jens Behrmann;Nicholas Carlini;Nicolas Papernot;J. Jacobsen

文献摘要

被引文献

相似文献

对抗性示例是恶意输入,旨在诱导错误分类。通常研究的基于敏感性的对抗性示例会对输入进行语义上的微小更改,从而导致不同的模型预测。本文研究了一种互补的失败模式,基于不变性的对抗性示例,它引入了最小的语义变化,修改了输入的真实标签,但保留了模型的预测。我们展示了这两种类型的对抗性例子之间的基本权衡。我们发现,防御基于敏感性的攻击,积极损害模型的准确性不变性为基础的攻击,并需要新的方法来抵御这两种攻击类型。特别是,我们通过生成模型(可证明)鲁棒的小扰动来打破最先进的经过对抗训练且经过验证的鲁棒模型,但这些小扰动会根据人类标签者改变输入的类别。最后,我们正式表明,过度不变的分类器的存在产生于标准数据集中过度强大的预测功能的存在。
Adversarial examples are malicious inputs crafted to induce misclassification. Commonly studied sensitivity-based adversarial examples introduce semantically-small changes to an input that result in a different model prediction. This paper studies a complementary failure mode, invariance-based adversarial examples, that introduce minimal semantic changes that modify an input's true label yet preserve the model's prediction. We demonstrate fundamental tradeoffs between these two types of adversarial examples. We show that defenses against sensitivity-based attacks actively harm a model's accuracy on invariance-based attacks, and that new approaches are needed to resist both attack types. In particular, we break state-of-the-art adversarially-trained and certifiably-robust models by generating small perturbations that the models are (provably) robust to, yet that change an input's class according to human labelers. Finally, we formally show that the existence of excessively invariant classifiers arises from the presence of overly-robust predictive features in standard datasets.