Toward Efficient and Adaptive Design of Video Detection System with Deep Neural Networks

Toward Efficient and Adaptive Design of Video Detection System with Deep Neural Networks
复制标题

DOI:
10.1145/3484946
复制
发表时间:
2022-05
期刊:
ACM Transactions on Embedded Computing Systems (TECS)
影响因子:
--
通讯作者:
Jiachen Mao;Qing Yang;Ang Li;Kent W. Nixon;H. Li;Yiran Chen
Jiachen Mao;Qing Yang;Ang Li;Kent W. Nixon;H. Li;Yiran Chen
中科院分区:
其他
文献类型:
--
作者:
Jiachen Mao;Qing Yang;Ang Li;Kent W. Nixon;H. Li;Yiran Chen

文献摘要

相似文献

在过去的十年中,深度神经网络(DNN),例如卷积神经网络,在目标分类和检测等视觉任务中实现了人类水平的性能。然而,众所周知,DNN 的计算成本很高,因此很难部署在实时和边缘应用中。之前的许多工作都集中在 DNN 模型压缩上,以获得更小的参数大小,从而降低计算成本。然而,此类方法通常会导致精度明显下降。在这项工作中,我们使用三个提出的想法优化了最先进的基于 DNN 的视频检测框架——来自云端的深度特征流 (DFF)。首先,我们提出异步 DFF (ADFF) 来异步执行神经网络。其次,我们提出了一种基于视频的动态调度(VDS)方法,该方法根据视频帧之间的运动幅度来决定检测频率。最后,我们提出了空间稀疏推理,它仅对部分视频帧进行推理,从而降低了计算成本。根据我们的实验结果,ADFF 可以将瓶颈延迟从 89 毫秒减少到 19 毫秒。 VDS 在不增加计算成本的情况下将检测精度提高了 0.6% mAP。 SSI 进一步节省了 0.2 毫秒,但检测精度 mAP 下降了 0.6%。
In the past decade, Deep Neural Networks (DNNs), e.g., Convolutional Neural Networks, achieved human-level performance in vision tasks such as object classification and detection. However, DNNs are known to be computationally expensive and thus hard to be deployed in real-time and edge applications. Many previous works have focused on DNN model compression to obtain smaller parameter sizes and consequently, less computational cost. Such methods, however, often introduce noticeable accuracy degradation. In this work, we optimize a state-of-the-art DNN-based video detection framework—Deep Feature Flow (DFF) from the cloud end using three proposed ideas. First, we propose Asynchronous DFF (ADFF) to asynchronously execute the neural networks. Second, we propose a Video-based Dynamic Scheduling (VDS) method that decides the detection frequency based on the magnitude of movement between video frames. Last, we propose Spatial Sparsity Inference, which only performs the inference on part of the video frame and thus reduces the computation cost. According to our experimental results, ADFF can reduce the bottleneck latency from 89 to 19 ms. VDS increases the detection accuracy by 0.6% mAP without increasing computation cost. And SSI further saves 0.2 ms with a 0.6% mAP degradation of detection accuracy.