Video-zilla: An Indexing Layer for Large-Scale Video Analytics

Video-zilla: An Indexing Layer for Large-Scale Video Analytics
复制标题

DOI:
10.1145/3514221.3517840
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 International Conference on Management of Data
影响因子:
--
通讯作者:
Bo Hu;Peizhen Guo;Wenjun Hu
Bo Hu;Peizhen Guo;Wenjun Hu
中科院分区:
其他
文献类型:
--
作者:
Bo Hu;Peizhen Guo;Wenjun Hu

文献摘要

相似文献

如今,监控摄像机的普遍部署对在许多摄像机源上运行的视频分析系统提出了巨大的可扩展性挑战。目前,除了标准文件系统提供的索引工具外,几乎没有索引工具来组织视频源。最近的视频分析系统实现了应用特定的帧分析和采样技术,以减少处理的原始视频的数量,利用帧级冗余或手动标记的相机之间的时空相关性。本文介绍了Video-zilla,一个独立的索引层之间的视频查询系统和视频存储来组织视频数据。我们提出了一个视频数据单元的抽象,语义视频流(SVS),基于视频中的对象之间的距离的概念。SVS隐式地捕获场景,这是当前视频内容表征中缺少的,也是单个帧和整个摄像机馈送之间的中间地带。然后,我们构建一个分层索引,暴露内部和跨摄像头提要的语义相似性,这样Video-zilla就可以根据视频提要的内容语义快速聚类视频提要,而无需手动标记。我们在三个用例中实现和评估Video-zilla:对象识别查询,用于训练专用DNN的聚类和存档服务。在所有这三种情况下,Video-zilla将摄像机间视频分析的时间复杂度从与摄像机数量的线性关系降低到次线性关系,并将查询资源使用量减少了14倍,而不是使用现有查询系统中内置的帧级或时空相似性。
Pervasive deployment of surveillance cameras today poses enormous scalability challenges to video analytics systems operating over many camera feeds. Currently, there are few indexing tools to organize video feeds beyond what is provided by a standard file system. Recent video analytic systems implement application-specific frame profiling and sampling techniques to reduce the number of raw videos processed, leveraging frame-level redundancy or manually labeled spatial-temporal correlation between cameras. This paper presents Video-zilla, a standalone indexing layer between video query systems and a video store to organize video data. We propose a video data unit abstraction, semantic video stream (SVS), based on a notion of distance between objects in the video. SVS implicitly captures scenes, which is missing from current video content characterization and a middle ground between individual frames and an entire camera feed. We then build a hierarchical index that exposes the semantic similarity both within and across camera feeds, such that Video-zilla can quickly cluster video feeds based on their content semantics without manual labeling. We implement and evaluate Video-zilla in three use cases: object identification queries, clustering for training specialized DNNs, and archival services. In all three cases, Video-zilla reduces the time complexity of inter-camera video analytics from linear with the number of cameras to sublinear, and reduces query resource usage by up to 14× compared to using frame-level or spatial-temporal similarity built into existing query systems.