Pack and Detect: Fast Object Detection in Videos Using Region-of-Interest Packing

Pack and Detect: Fast Object Detection in Videos Using Region-of-Interest Packing
复制标题

DOI:
10.1145/3297001.3297020
复制
发表时间:
2018-09
期刊:
Proceedings of the ACM India Joint International Conference on Data Science and Management of Data
影响因子:
--
通讯作者:
Athindran Ramesh Kumar;Balaraman Ravindran;A. Raghunathan
Athindran Ramesh Kumar;Balaraman Ravindran;A. Raghunathan
中科院分区:
其他
文献类型:
--
作者:
Athindran Ramesh Kumar;Balaraman Ravindran;A. Raghunathan

文献摘要

被引文献

相似文献

视频中的目标检测是计算机视觉中的一项重要任务,用于目标跟踪、视频摘要和视频搜索等各种应用。尽管近年来由于深度神经网络的兴起,在提高目标检测的准确性方面取得了很大进展,但最先进的算法是高度计算密集型的。为了应对这一挑战,我们在视频的上下文中进行了两个重要的观察:(i)对象通常只占用每个视频帧中的一小部分区域,以及(ii)连续帧之间存在强时间相关性的可能性很高。基于这些观察,我们提出了包装和检测(PaD),一种方法,以减少视频中的对象检测的计算要求。在PaD中,仅以全尺寸处理被称为锚帧的选定视频帧。在位于锚帧之间的帧(锚帧间)中,基于前一帧中的检测来识别感兴趣区域(ROI)。我们提出了一种算法,将每个锚点间帧的ROI打包到一个尺寸减小的帧中。检测器的计算要求由于输入的较小尺寸而降低。为了保持目标检测的准确性,该算法对感兴趣区域进行了扩展,为检测器提供了每个目标周围的额外背景。PaD可以使用任何底层神经网络架构来处理全尺寸和缩小尺寸的帧。使用ImageNet视频对象检测数据集的实验表明,PaD可以将一帧所需的FLOPS数量减少4倍。这导致在配备NVIDIA Titan X GPU的2.1 GHz Intel Xeon服务器上的吞吐量总体增加1.25倍,但精度下降1.1%。
Object detection in videos is an important task in computer vision for various applications such as object tracking, video summarization and video search. Although great progress has been made in improving the accuracy of object detection in recent years due to the rise of deep neural networks, the state-of-the-art algorithms are highly computationally intensive. In order to address this challenge, we make two important observations in the context of videos: (i) Objects often occupy only a small fraction of the area in each video frame, and (ii) There is a high likelihood of strong temporal correlation between consecutive frames. Based on these observations, we propose Pack and Detect (PaD), an approach to reduce the computational requirements of object detection in videos. In PaD, only selected video frames called anchor frames are processed at full size. In the frames that lie between anchor frames (inter-anchor frames), regions of interest (ROIs) are identified based on the detections in the previous frame. We propose an algorithm to pack the ROIs of each inter-anchor frame together into a reduced-size frame. The computational requirements of the detector are reduced due to the lower size of the input. In order to maintain the accuracy of object detection, the proposed algorithm expands the ROIs greedily to provide additional background around each object to the detector. PaD can use any underlying neural network architecture to process the full-size and reduced-size frames. Experiments using the ImageNet video object detection dataset indicate that PaD can potentially reduce the number of FLOPS required for a frame by 4×. This leads to an overall increase in throughput of 1.25× on a 2.1 GHz Intel Xeon server with a NVIDIA Titan X GPU at the cost of 1.1% drop in accuracy.