Instance, Scale, and Teacher Adaptive Knowledge Distillation for Visual Detection in Autonomous Driving

Instance, Scale, and Teacher Adaptive Knowledge Distillation for Visual Detection in Autonomous Driving
复制标题

DOI:
10.1109/tiv.2022.3217261
复制
发表时间:
2023-03-01
影响因子:
8.2
通讯作者:
Tian, Qing
Tian, Qing
中科院分区:
工程技术2区
文献类型:
--
作者:
Lan, Qizhen;Tian, Qing

文献摘要

被引文献

相似文献

高效的视觉检测是自动驾驶感知的关键组成部分,为后期规划和控制阶段奠定基础。基于深度网络的视觉系统实现了最先进的性能,但对于嵌入式设备(例如行车记录仪)来说,它们通常很麻烦且计算上不可行。知识蒸馏是推导更高效模型的有效方法。然而,大多数现有的工作都以分类任务为目标,并平等地对待所有实例。在本文中,我们首先提出了用于自动驾驶视觉检测的自适应实例蒸馏(AID)方法。它可以根据教师的损失,通过重新权衡每个实例和每个尺度进行蒸馏,选择性地将教师的知识传授给学生。此外,为了使学生能够有效地消化来自多个来源的知识,我们还提出了多教师自适应实例蒸馏(M-AID)方法。我们的 M-AID 帮助学生从每位老师那里学习最好的知识。某些实例和规模。与之前的 KD 方法不同,我们的 M-AID 以实例、规模和教师自适应的方式调整蒸馏权重。在 KITTI、COCO-Traffic 和 SODA10 M 数据集上的实验表明,我们的方法提高了自动驾驶场景中不同检测器上各种最先进的 KD 方法的性能。与基线相比,我们的 AID 使单级和两级探测器的 mAP 分别平均增加 2.28% 和 2.98%。通过战略性地整合多位教师的知识,我们的 M-AID 方法平均提高了 2.92% 的 mAP。
Efficient visual detection is a crucial component in self-driving perception and lays the foundation for later planning and control stages. Deep-networks-based visual systems achieve state-of-the-art performance, but they are usually cumbersome and computationally infeasible for embedded devices (e.g., dash cams). Knowledge distillation is an effective way to derive more efficient models. However, most existing works target classification tasks and treat all instances equally. In this paper, we first present our Adaptive Instance Distillation (AID) method for self-driving visual detection. It can selectively impart the teacher's knowledge to the student by re-weighing each instance and each scale for distillation based on the teacher's loss. In addition, to enable the student to effectively digest knowledge from multiple sources, we also propose a Multi-Teacher Adaptive Instance Distillation (M-AID) method. Our M-AID helps the student to learn the best knowledge from each teacher w.r.t. certain instances and scales. Unlike previous KD methods, our M-AID adjusts the distillation weights in an instance, scale, and teacher adaptive manner. Experiments on the KITTI, COCO-Traffic, and SODA10 M datasets show that our methods improve the performance of a wide variety of state-of-the-art KD methods on different detectors in self-driving scenarios. Compared to the baseline, our AID leads to an average of 2.28% and 2.98% mAP increases for single-stage and two-stage detectors, respectively. By strategically integrating knowledge from multiple teachers, our M-AID method achieves an average of 2.92% mAP improvement.