Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks

Gotta Catch'Em All: Using Honeypots to Catch Adversarial Attacks on Neural Networks
复制标题

DOI:
10.1145/3372297.3417231
复制
发表时间:
2019-04
期刊:
Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Shawn Shan;Emily Wenger;Bolun Wang;B. Li;Haitao Zheng;Ben Y. Zhao
Shawn Shan;Emily Wenger;Bolun Wang;B. Li;Haitao Zheng;Ben Y. Zhao
中科院分区:
其他
文献类型:
--
作者:
Shawn Shan;Emily Wenger;Bolun Wang;B. Li;Haitao Zheng;Ben Y. Zhao

文献摘要

被引文献

相似文献

深度神经网络(DNN)容易受到对抗性攻击。许多努力要么试图修补训练模型中的弱点,要么试图使计算利用它们的对抗性示例变得困难或昂贵。在我们的工作中,我们探索了一种新的“蜜罐”方法来保护DNN模型。我们故意在分类流形中注入陷阱门,蜜罐弱点,吸引攻击者搜索对抗性的例子。攻击者的优化算法倾向于陷阱门,导致他们在特征空间中产生类似于陷阱门的攻击。然后,我们的防御通过将输入的神经元激活签名与陷阱门的神经元激活签名进行比较来识别攻击。在本文中,我们介绍了陷门,并描述了一个陷门启用防御的实现。首先,我们分析证明了陷阱门塑造了对抗性攻击的计算,使得攻击输入将具有与陷阱门非常相似的特征表示。其次,我们通过实验证明,陷阱保护模型可以高精度地检测到由最先进的攻击(PGD,基于优化的CW,弹性网络,BPDA)生成的对抗性示例,对正常分类的影响可以忽略不计。这些结果概括了分类领域,包括图像,面部和交通标志识别。我们还提出了显着的结果测量陷阱的鲁棒性对定制的自适应攻击(对策)。
Deep neural networks (DNN) are known to be vulnerable to adversarial attacks. Numerous efforts either try to patch weaknesses in trained models, or try to make it difficult or costly to compute adversarial examples that exploit them. In our work, we explore a new "honeypot" approach to protect DNN models. We intentionally inject trapdoors, honeypot weaknesses in the classification manifold that attract attackers searching for adversarial examples. Attackers' optimization algorithms gravitate towards trapdoors, leading them to produce attacks similar to trapdoors in the feature space. Our defense then identifies attacks by comparing neuron activation signatures of inputs to those of trapdoors. In this paper, we introduce trapdoors and describe an implementation of a trapdoor-enabled defense. First, we analytically prove that trapdoors shape the computation of adversarial attacks so that attack inputs will have feature representations very similar to those of trapdoors. Second, we experimentally show that trapdoor-protected models can detect, with high accuracy, adversarial examples generated by state-of-the-art attacks (PGD, optimization-based CW, Elastic Net, BPDA), with negligible impact on normal classification. These results generalize across classification domains, including image, facial, and traffic-sign recognition. We also present significant results measuring trapdoors' robustness against customized adaptive attacks (countermeasures).