Why is the video analytics accuracy fluctuating, and what can we do about it?

Why is the video analytics accuracy fluctuating, and what can we do about it?
复制标题

DOI:
10.48550/arxiv.2208.12644
复制
发表时间:
2022-08
期刊:
--
影响因子:
--
通讯作者:
Sibendu Paul;Kunal Rao;G. Coviello;Murugan Sankaradas;Oliver Po;Y. C. Hu;S. Chakradhar
Sibendu Paul;Kunal Rao;G. Coviello;Murugan Sankaradas;Oliver Po;Y. C. Hu;S. Chakradhar
中科院分区:
其他
文献类型:
--
作者:
Sibendu Paul;Kunal Rao;G. Coviello;Murugan Sankaradas;Oliver Po;Y. C. Hu;S. Chakradhar

文献摘要

相似文献

将视频视为一系列图像(帧)是一种常见的做法,并重复使用仅在图像上训练的深度神经网络模型来完成类似的视频分析任务。在本文中,我们展示了这种信念的飞跃,即在图像上工作良好的深度学习模型也可以在视频上工作良好,实际上是有缺陷的。我们表明,即使摄像机正在观看一个没有以任何人类可感知的方式变化的场景,并且我们控制了视频压缩和环境(照明)等外部因素,视频分析应用程序的准确性也会明显波动。这些波动的发生是因为视频摄像机产生的连续帧可能在视觉上看起来相似,但视频分析应用程序对这些帧的感知却大不相同。我们观察到,这些波动的根本原因是摄像机为了捕捉和产生视觉上令人愉悦的视频而自动进行的动态摄像机参数变化。摄像机无意中充当了一个无意的对手,因为在连续帧中图像像素值的这些微小变化,正如我们所示,对重复使用图像训练的深度学习模型的视频分析任务的见解的准确性有明显的不利影响。为了解决这种来自摄像机的无意的对抗效应,我们探索使用迁移学习技术,通过从图像分析任务中学习的知识转移来提高视频分析任务中的学习。特别是,我们表明我们新训练的Yolov5模型减少了跨帧对象检测的波动,从而更好地跟踪对象(跟踪错误减少40%)。我们的论文还提供了新的方向和技术,以减轻摄像机对用于视频分析应用的深度学习模型的对抗效应。
It is a common practice to think of a video as a sequence of images (frames), and re-use deep neural network models that are trained only on images for similar analytics tasks on videos. In this paper, we show that this leap of faith that deep learning models that work well on images will also work well on videos is actually flawed. We show that even when a video camera is viewing a scene that is not changing in any human-perceptible way, and we control for external factors like video compression and environment (lighting), the accuracy of video analytics application fluctuates noticeably. These fluctuations occur because successive frames produced by the video camera may look similar visually, but these frames are perceived quite differently by the video analytics applications. We observed that the root cause for these fluctuations is the dynamic camera parameter changes that a video camera automatically makes in order to capture and produce a visually pleasing video. The camera inadvertently acts as an unintentional adversary because these slight changes in the image pixel values in consecutive frames, as we show, have a noticeably adverse impact on the accuracy of insights from video analytics tasks that re-use image-trained deep learning models. To address this inadvertent adversarial effect from the camera, we explore the use of transfer learning techniques to improve learning in video analytics tasks through the transfer of knowledge from learning on image analytics tasks. In particular, we show that our newly trained Yolov5 model reduces fluctuation in object detection across frames, which leads to better tracking of objects(40% fewer mistakes in tracking). Our paper also provides new directions and techniques to mitigate the camera's adversarial effect on deep learning models used for video analytics applications.