Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object Tracking

Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object Tracking
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
Yunhan Jia;Yantao Lu;Junjie Shen;Qi Alfred Chen;Hao Chen;Zhenyu Zhong;Tao Wei
Yunhan Jia;Yantao Lu;Junjie Shen;Qi Alfred Chen;Hao Chen;Zhenyu Zhong;Tao Wei
中科院分区:
其他
文献类型:
--
作者:
Yunhan Jia;Yantao Lu;Junjie Shen;Qi Alfred Chen;Hao Chen;Zhenyu Zhong;Tao Wei

文献摘要

被引文献

相似文献

对抗机器学习的最新研究开始关注自动驾驶中的视觉感知,并研究了对象检测模型的对抗示例(AE)。然而,在这种视觉感知管道中,还必须在称为多对象跟踪(MOT)的过程中跟踪检测到的对象,以构建周围障碍物的移动轨迹。由于MOT被设计为对目标检测中的错误具有鲁棒性,因此它对盲目针对目标检测的现有攻击技术提出了普遍挑战:我们发现,需要超过98%的成功率才能真正影响跟踪结果,这是现有攻击技术无法满足的要求。在本文中,我们首次研究了对抗机器学习攻击自动驾驶中完整的视觉感知管道,并发现了一种新的攻击技术,跟踪器劫持,可以有效地欺骗MOT使用AE进行对象检测。使用我们的技术,只要在一个帧上成功地进行AE,就可以将现有物体移入或移出自动驾驶汽车的车头间距,从而造成潜在的安全隐患。我们使用Berkeley Deep Drive数据集进行评估,发现平均攻击3帧时,我们的攻击可以有近100%的成功率,而盲目目标检测的攻击只有25%。
Recent work in adversarial machine learning started to focus on the visual perception in autonomous driving and studied Adversarial Examples (AEs) for object detection models. However, in such visual perception pipeline the detected objects must also be tracked, in a process called Multiple Object Tracking (MOT), to build the moving trajectories of surrounding obstacles. Since MOT is designed to be robust against errors in object detection, it poses a general challenge to existing attack techniques that blindly target objection detection: we find that a success rate of over 98% is needed for them to actually affect the tracking results, a requirement that no existing attack technique can satisfy. In this paper, we are the first to study adversarial machine learning attacks against the complete visual perception pipeline in autonomous driving, and discover a novel attack technique, tracker hijacking, that can effectively fool MOT using AEs on object detection. Using our technique, successful AEs on as few as one single frame can move an existing object in to or out of the headway of an autonomous vehicle to cause potential safety hazards. We perform evaluation using the Berkeley Deep Drive dataset and find that on average when 3 frames are attacked, our attack can have a nearly 100% success rate while attacks that blindly target object detection only have up to 25%.