Global Memory and Local Continuity for Video Object Detection

Global Memory and Local Continuity for Video Object Detection
复制标题

DOI:
10.1109/tmm.2022.3164253
复制
发表时间:
2023
影响因子:
7.3
通讯作者:
Liang Han;Zhaozheng Yin
Liang Han;Zhaozheng Yin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liang Han;Zhaozheng Yin

文献摘要

相似文献

为了应对视频对象检测(VOD)中的挑战,例如遮挡和运动模糊,许多最先进的视频对象检测器采用特征聚合模块来编码远程上下文信息以支持当前帧。这些检测器的主要缺点有三个:首先,逐帧检测降低了检测速度;其次,逐帧检测通常会忽略视频中对象的局部连续性,导致检测时间不一致;第三,特征聚合模块通常对本地视频剪辑或单个视频中的时间特征进行编码,而不利用其他视频中的特征。在这项工作中,我们开发了一种在线 VOD 算法,通过利用全局内存和局部连续性,实现高速和高精度的平衡。在算法中,设计了一个有效且高效的全局存储库(GMB)来存储和更新对象类特征,这使我们能够利用其他视频中的支持特征来增强当前视频帧中的对象特征。此外,为了进一步加快检测速度,我们设计了一个对象跟踪器,利用视频的局部连续性特性,根据关键帧的检测结果对非关键帧进行对象检测。考虑到检测精度和速度之间的权衡,所提出的框架在 ImageNet VID 数据集上实现了卓越的性能。源代码将通过我们的 GitHub 网站向公众发布。
To deal with the challenges in video object detection (VOD), such as occlusion and motion blur, many state-of-the-art video object detectors adopt a feature aggregation module to encode the long-range contextual information to support the current frame. The main drawbacks of these detectors are three-folds: first, the frame-wise detection slows down the detection speed; second, the frame-wise detection usually ignores the local continuity of the objects in a video, resulting in temporal inconsistent detection; third, the feature aggregation module usually encodes temporal features either from a local video clip or a single video, without exploiting the features in other videos. In this work, we develop an online VOD algorithm, aiming at a balanced high-speed and high-accuracy, by exploiting the global memory and local continuity. In the algorithm, an effective and efficient global memory bank (GMB) is designed to deposit and update object class features, which enables us to exploit the support features in other videos to enhance object features in the current video frames. Besides, to further speed up the detection, we design an object tracker to perform object detection for non-key frames based on the detection results of the key frame by leveraging the local continuity property of the video. Considering the trade-off between detection accuracy and speed, the proposed framework achieves superior performance on the ImageNet VID dataset. Source codes will be released to the public via our GitHub website.