Surveillance Video Parsing with Single Frame Supervision

Surveillance Video Parsing with Single Frame Supervision
复制标题

DOI:
10.1109/cvpr.2017.114
复制
发表时间:
2016-11
期刊:
2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Si Liu;Changhu Wang;Ruihe Qian;Han Yu;Renda Bao;Yao Sun
Si Liu;Changhu Wang;Ruihe Qian;Han Yu;Renda Bao;Yao Sun
中科院分区:
其他
文献类型:
--
作者:
Si Liu;Changhu Wang;Ruihe Qian;Han Yu;Renda Bao;Yao Sun

文献摘要

被引文献

相似文献

监控视频解析将视频帧分割为几个标签,如脸、裤子、左腿,具有广泛的应用[41,8]。然而,明智地对所有帧进行像素注释是乏味和低效的。在本文中,我们开发了一种单帧视频解析(SVP)方法,该方法在训练阶段每个视频只需要一个标记帧。为了解析一个特定的帧,将联合考虑该帧之前的视频片段。SVP (i)对视频片段内的帧进行粗略解析,(ii)对帧之间的光流进行估计,(iii)对被光流扭曲的粗略解析结果进行融合,得到精细化的解析结果。将帧解析、光流估计和时间融合三个部分端到端集成在一起。在两个监控视频数据集上的实验结果表明了SVP算法的优越性。收集到的视频解析数据集可以通过http://liusi-group.com/projects/SVP下载,以供进一步研究。
Surveillance video parsing, which segments the video frames into several labels, e.g., face, pants, left-leg, has wide applications [41, 8]. However, pixel-wisely annotating all frames is tedious and inefficient. In this paper, we develop a Single frame Video Parsing (SVP) method which requires only one labeled frame per video in training stage. To parse one particular frame, the video segment preceding the frame is jointly considered. SVP (i) roughly parses the frames within the video segment, (ii) estimates the optical flow between frames and (iii) fuses the rough parsing results warped by optical flow to produce the refined parsing result. The three components of SVP, namely frame parsing, optical flow estimation and temporal fusion are integrated in an end-to-end manner. Experimental results on two surveillance video datasets show the superiority of SVP over state-of-the-arts. The collected video parsing datasets can be downloaded via http://liusi-group.com/projects/SVP for the further studies.