AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling

AdaScale: Towards Real-time Video Object Detection Using Adaptive Scaling
复制标题

DOI:
--
复制
发表时间:
2019-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Ting-Wu Chin;Ruizhou Ding;Diana Marculescu
Ting-Wu Chin;Ruizhou Ding;Diana Marculescu
中科院分区:
其他
文献类型:
--
作者:
Ting-Wu Chin;Ruizhou Ding;Diana Marculescu

文献摘要

相似文献

在机器人和自动驾驶汽车等视觉自主系统中,视频目标检测起着至关重要的作用,其速度和准确性是提供可靠运行的重要因素。我们在本文中展示的关键见解是,当涉及到图像缩放时,速度和准确性不一定是一个权衡。我们的结果表明,将图像重新缩放到较低的分辨率有时会产生更好的精度。基于这一观察,我们提出了一种新的方法,称为AdaScale,它自适应地选择输入图像尺度,提高了视频目标检测的准确性和速度。为此,我们在ImageNet VID和mini YouTube-BoundingBoxes数据集上的结果显示,mAP分别提高了1.3点和2.7点,加速速度分别提高了1.6倍和1.8倍。此外,我们通过在ImageNet VID数据集上稍微更好的mAP,将最先进的视频加速工作提高了1.25倍。
In vision-enabled autonomous systems such as robots and autonomous cars, video object detection plays a crucial role, and both its speed and accuracy are important factors to provide reliable operation. The key insight we show in this paper is that speed and accuracy are not necessarily a trade-off when it comes to image scaling. Our results show that re-scaling the image to a lower resolution will sometimes produce better accuracy. Based on this observation, we propose a novel approach, dubbed AdaScale, which adaptively selects the input image scale that improves both accuracy and speed for video object detection. To this end, our results on ImageNet VID and mini YouTube-BoundingBoxes datasets demonstrate 1.3 points and 2.7 points mAP improvement with 1.6x and 1.8x speedup, respectively. Additionally, we improve state-of-the-art video acceleration work by an extra 1.25x speedup with slightly better mAP on ImageNet VID dataset.