RILOD: near real-time incremental learning for object detection at the edge

RILOD: near real-time incremental learning for object detection at the edge
复制标题

DOI:
10.1145/3318216.3363317
复制
发表时间:
2019-03
期刊:
Proceedings of the 4th ACM/IEEE Symposium on Edge Computing
影响因子:
--
通讯作者:
Dawei Li;Serafettin Tasci;Shalini Ghosh-;Jingwen Zhu;Junting Zhang;Larry Heck
Dawei Li;Serafettin Tasci;Shalini Ghosh-;Jingwen Zhu;Junting Zhang;Larry Heck
中科院分区:
其他
文献类型:
--
作者:
Dawei Li;Serafettin Tasci;Shalini Ghosh-;Jingwen Zhu;Junting Zhang;Larry Heck

文献摘要

被引文献

相似文献

配备摄像头的边缘设备附带的对象检测模型无法覆盖每个用户感兴趣的对象。因此,增量学习能力是一个强大的和个性化的对象检测系统,许多应用程序将依赖于一个重要的功能,在本文中,我们提出了一个高效而实用的系统,RILOD,增量训练现有的对象检测模型,使其可以检测新的对象类,而不会失去其检测旧类的能力。RILOD的关键组件是一种新型的增量学习算法,该算法仅使用新对象类的训练数据来训练一阶段深度对象检测模型的端到端。具体来说,为了避免灾难性的遗忘,该算法从旧模型中提取三种类型的知识,以模仿旧模型在对象分类、边界框回归和特征提取方面的行为。此外,由于新类的训练数据可能不可用,因此实时数据集构建管道被设计为实时收集训练图像,并使用类别和边界框注释自动标记图像。我们已经在边缘云和仅边缘设置下实现了RILOD。实验结果表明,该系统可以在几分钟内学会检测一个新的对象类,包括数据集的构建和模型训练。相比之下,传统的基于微调的方法可能需要几个小时的训练,并且在大多数情况下还需要繁琐且昂贵的手动数据集标记步骤。
Object detection models shipped with camera-equipped edge devices cannot cover the objects of interest for every user. Therefore, the incremental learning capability is a critical feature for a robust and personalized object detection system that many applications would rely on. In this paper, we present an efficient yet practical system, RILOD, to incrementally train an existing object detection model such that it can detect new object classes without losing its capability to detect old classes. The key component of RILOD is a novel incremental learning algorithm that trains end-to-end for one-stage deep object detection models only using training data of new object classes. Specifically to avoid catastrophic forgetting, the algorithm distills three types of knowledge from the old model to mimic the old model's behavior on object classification, bounding box regression and feature extraction. In addition, since the training data for the new classes may not be available, a real-time dataset construction pipeline is designed to collect training images on-the-fly and automatically label the images with both category and bounding box annotations. We have implemented RILOD under both edge-cloud and edge-only setups. Experiment results show that the proposed system can learn to detect a new object class in just a few minutes, including both dataset construction and model training. In comparison, traditional fine-tuning based method may take a few hours for training, and in most cases would also need a tedious and costly manual dataset labeling step.