Physical Adversarial Examples for Object Detectors

Physical Adversarial Examples for Object Detectors
复制标题

DOI:
--
复制
发表时间:
2018-07
期刊:
ArXiv
影响因子:
--
通讯作者:
Kevin Eykholt;I. Evtimov;Earlence Fernandes;Bo Li;Amir Rahmati;Florian Tramèr;Atul Prakash;
Kevin Eykholt;I. Evtimov;Earlence Fernandes;Bo Li;Amir Rahmati;Florian Tramèr;Atul Prakash;
中科院分区:
其他
文献类型:
--
作者:
Kevin Eykholt;I. Evtimov;Earlence Fernandes;Bo Li;Amir Rahmati;Florian Tramèr;Atul Prakash;

文献摘要

被引文献

相似文献

深度神经网络(DNN)容易受到对抗样本的攻击——这些恶意构造的输入会导致DNN做出错误的预测。近期的研究表明,这些攻击可推广到物理领域,在各种现实条件下对物理对象制造干扰,从而欺骗图像分类器。此类攻击对安全关键型信息物理系统中使用的深度学习模型构成了风险。在这项工作中,我们将物理攻击扩展到更具挑战性的目标检测模型,这是一类更广泛的深度学习算法,广泛用于检测和标记场景中的多个目标。在对图像分类器先前的物理攻击基础上进行改进,我们制造出受干扰的物理对象,这些对象会被目标检测模型忽略或错误标记。我们实施了一种“消失攻击”,即我们让一个停止标志根据探测器“消失”——要么用对抗性的停止标志海报覆盖该标志,要么在标志上添加对抗性贴纸。在一个受控实验室环境中录制的视频里,最先进的YOLOv2探测器在超过85%的视频帧中未能识别出这些对抗性的停止标志。在一次户外实验中,YOLO分别在72.5%和63.5%的视频帧中被海报和贴纸攻击所欺骗。我们还使用了一种不同的目标检测模型Faster R - CNN来证明我们的对抗性干扰的可转移性。所制造的海报干扰能够在受控实验室环境中85.9%的视频帧以及户外环境中40.2%的视频帧中欺骗Faster R - CNN。最后,我们展示了一种新的“创造攻击”的初步结果,即无害的物理贴纸欺骗模型检测出不存在的目标。
Deep neural networks (DNNs) are vulnerable to adversarial examples-maliciously crafted inputs that cause DNNs to make incorrect predictions. Recent work has shown that these attacks generalize to the physical domain, to create perturbations on physical objects that fool image classifiers under a variety of real-world conditions. Such attacks pose a risk to deep learning models used in safety-critical cyber-physical systems. In this work, we extend physical attacks to more challenging object detection models, a broader class of deep learning algorithms widely used to detect and label multiple objects within a scene. Improving upon a previous physical attack on image classifiers, we create perturbed physical objects that are either ignored or mislabeled by object detection models. We implement a Disappearance Attack, in which we cause a Stop sign to "disappear" according to the detector-either by covering thesign with an adversarial Stop sign poster, or by adding adversarial stickers onto the sign. In a video recorded in a controlled lab environment, the state-of-the-art YOLOv2 detector failed to recognize these adversarial Stop signs in over 85% of the video frames. In an outdoor experiment, YOLO was fooled by the poster and sticker attacks in 72.5% and 63.5% of the video frames respectively. We also use Faster R-CNN, a different object detection model, to demonstrate the transferability of our adversarial perturbations. The created poster perturbation is able to fool Faster R-CNN in 85.9% of the video frames in a controlled lab environment, and 40.2% of the video frames in an outdoor environment. Finally, we present preliminary results with a new Creation Attack, where in innocuous physical stickers fool a model into detecting nonexistent objects.