ML-LOO: Detecting Adversarial Examples with Feature Attribution

ML-LOO: Detecting Adversarial Examples with Feature Attribution
复制标题

DOI:
10.1609/aaai.v34i04.6140
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Puyudi Yang;Jianbo Chen;Cho-Jui Hsieh;Jane-ling Wang;Michael I. Jordan
Puyudi Yang;Jianbo Chen;Cho-Jui Hsieh;Jane-ling Wang;Michael I. Jordan
中科院分区:
其他
文献类型:
--
作者:
Puyudi Yang;Jianbo Chen;Cho-Jui Hsieh;Jane-ling Wang;Michael I. Jordan

文献摘要

被引文献

相似文献

深度神经网络在一系列任务上获得了最先进的性能。然而,他们很容易通过向输入添加一个小的对抗性扰动来欺骗。在图像数据上的扰动通常是人类无法察觉的。我们观察到一个显着的差异,功能之间的属性adversarially制作的例子和原始的例子。基于这一观察结果,我们引入了一个新的框架,通过对特征归因得分的尺度估计进行阈值化来检测对抗性示例。此外,我们扩展了我们的方法,包括多层特征属性,以解决具有混合置信度的攻击。如在广泛的实验中所示,我们的方法实现了上级性能区分对抗性的例子从流行的攻击方法在各种真实的数据集相比,国家的最先进的检测方法。特别是,我们的方法能够检测混合置信水平的对抗性示例,并在不同的攻击方法之间进行转换。我们还表明,我们的方法实现了竞争力的性能,即使当攻击者完全访问检测器。
Deep neural networks obtain state-of-the-art performance on a series of tasks. However, they are easily fooled by adding a small adversarial perturbation to the input. The perturbation is often imperceptible to humans on image data. We observe a significant difference in feature attributions between adversarially crafted examples and original examples. Based on this observation, we introduce a new framework to detect adversarial examples through thresholding a scale estimate of feature attribution scores. Furthermore, we extend our method to include multi-layer feature attributions in order to tackle attacks that have mixed confidence levels. As demonstrated in extensive experiments, our method achieves superior performances in distinguishing adversarial examples from popular attack methods on a variety of real data sets compared to state-of-the-art detection methods. In particular, our method is able to detect adversarial examples of mixed confidence levels, and transfer between different attacking methods. We also show that our method achieves competitive performance even when the attacker has complete access to the detector.