One Man’s Trash Is Another Man’s Treasure: Resisting Adversarial Examples by Adversarial Examples

One Man’s Trash Is Another Man’s Treasure: Resisting Adversarial Examples by Adversarial Examples
复制标题

DOI:
10.1109/cvpr42600.2020.00049
复制
发表时间:
2019-11
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Chang Xiao;Changxi Zheng
Chang Xiao;Changxi Zheng
中科院分区:
其他
文献类型:
--
作者:
Chang Xiao;Changxi Zheng

文献摘要

相似文献

现代图像分类系统通常建立在深度神经网络的基础上,该网络面临着对抗性例子——带有故意制作的、难以察觉的噪声的图像,会误导网络的分类。为了防御对抗性示例,一个可行的想法是混淆网络相对于输入图像的梯度。这个总体想法激发了一系列防御方法。然而,几乎所有这些都被证明是脆弱的。我们从一个完全不同的角度重新审视这个看似有缺陷的想法。我们拥抱无处不在的对抗性例子和制作它们的数字程序,并将这种有害的攻击过程转变为有用的防御机制。我们的防御方法在概念上很简单:在输入输入图像进行分类之前,通过在预先训练的外部模型上查找对抗性示例来对其进行转换。我们针对各种可能的攻击评估我们的方法。在 CIFAR-10 和 Tiny ImageNet 数据集上,我们的方法比最先进的方法更加稳健。特别是,与对抗性训练相比,我们的方法提供了更低的训练成本和更强的鲁棒性。
Modern image classification systems are often built on deep neural networks, which suffer from adversarial examples--images with deliberately crafted, imperceptible noise to mislead the network's classification. To defend against adversarial examples, a plausible idea is to obfuscate the network's gradient with respect to the input image. This general idea has inspired a long line of defense methods. Yet, almost all of them have proven vulnerable. We revisit this seemingly flawed idea from a radically different perspective. We embrace the omnipresence of adversarial examples and the numerical procedure of crafting them, and turn this harmful attacking process into a useful defense mechanism. Our defense method is conceptually simple: before feeding an input image for classification, transform it by finding an adversarial example on a pre-trained external model. We evaluate our method against a wide range of possible attacks. On both CIFAR-10 and Tiny ImageNet datasets, our method is significantly more robust than state-of-the-art methods. Particularly, in comparison to adversarial training, our method offers lower training cost as well as stronger robustness.