Detecting Adversarial Samples Using Influence Functions and Nearest Neighbors

Detecting Adversarial Samples Using Influence Functions and Nearest Neighbors
复制标题

DOI:
10.1109/cvpr42600.2020.01446
复制
发表时间:
2019-09
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Gilad Cohen;G. Sapiro;R. Giryes
Gilad Cohen;G. Sapiro;R. Giryes
中科院分区:
其他
文献类型:
--
作者:
Gilad Cohen;G. Sapiro;R. Giryes

文献摘要

被引文献

相似文献

深度神经网络(dnn)因其易受对抗性攻击而臭名昭著,对抗性攻击是在其输入图像中添加的小扰动,以误导其预测。因此,对抗性样本的检测是鲁棒分类框架的基本要求。在这项工作中,我们提出了一种检测这种对抗性攻击的方法,该方法适用于任何预训练的神经网络分类器。我们使用影响函数来度量每个训练样本对验证集数据的影响。从影响分数中,我们找到了任何给定验证示例中最支持的训练样本。在DNN的激活层上拟合一个k近邻(k-NN)模型来搜索这些支持训练样本的排名。我们观察到,这些样本与正常输入的最近邻居高度相关,而对抗性输入的相关性要弱得多。我们使用k-NN的秩和距离训练了一个对抗性检测器,并表明它成功地区分了对抗性示例,在三个数据集的六种攻击方法上获得了最先进的结果。代码可从https://github.com/giladcohen/NNIF_adv_defense获得。
Deep neural networks (DNNs) are notorious for their vulnerability to adversarial attacks, which are small perturbations added to their input images to mislead their prediction. Detection of adversarial examples is, therefore, a fundamental requirement for robust classification frameworks. In this work, we present a method for detecting such adversarial attacks, which is suitable for any pre-trained neural network classifier. We use influence functions to measure the impact of every training sample on the validation set data. From the influence scores, we find the most supportive training samples for any given validation example. A k-nearest neighbor (k-NN) model fitted on the DNN's activation layers is employed to search for the ranking of these supporting training samples. We observe that these samples are highly correlated with the nearest neighbors of the normal inputs, while this correlation is much weaker for adversarial inputs. We train an adversarial detector using the k-NN ranks and distances and show that it successfully distinguishes adversarial examples, getting state-of-the-art results on six attack methods with three datasets. Code is available at https://github.com/giladcohen/NNIF_adv_defense.