Maze: A Cost-Efficient Video Deduplication System at Web-scale

Maze: A Cost-Efficient Video Deduplication System at Web-scale
复制标题

DOI:
10.1145/3503161.3548145
复制
发表时间:
2022-10
期刊:
Proceedings of the 30th ACM International Conference on Multimedia
影响因子:
--
通讯作者:
An Qin;Mengbai Xiao;Ben Huang;Xiaodong Zhang
An Qin;Mengbai Xiao;Ben Huang;Xiaodong Zhang
中科院分区:
其他
文献类型:
--
作者:
An Qin;Mengbai Xiao;Ben Huang;Xiaodong Zhang

文献摘要

相似文献

随着互联网视频的进步和主导服务,基于内容的视频重复数据删除系统成为Internet视频服务的必不可少的基础架构。但是,关于互联网的爆炸性增长的视频数据以多种方式挑战了系统设计和实施的可扩展性。 (1)尽管基于量化的索引技术对于大规模搜索视觉特征有效,但必须定期进行整个数据集的昂贵重新训练。 (2)视觉特征的高维矢量需要越来越大的SSD空间,从而使I/O性能降低。 (3)从互联网上爬行的视频是多种多样的,视觉上相似的视频不一定是重复项,从而提高了重复数据删除的复杂性。 (4)大多数视频都是编辑的。重复的内容更有可能是视频中的剪辑,要求处理细节的处理技术。为了解决上述问题,我们提出了一个全面的视频删除系统迷宫。迷宫的ANN层索引和搜索高维特征向量。 ANN层的体系结构支持有效的读取和写入并消除由重新训练引起的数据迁移。迷宫采用基于CNN的功能和ORB功能作为视觉功能,可针对特定的视频重复数据删除任务进行了优化。这些功能是紧凑的,并且完全驻留在内存中。声学特征还包含在迷宫中,因此视觉上相似的视频但具有不同的音轨是可识别的。开发了一种基于夹子的匹配算法,以发现精细粒度的重复内容。迷宫已作为生产系统部署了两年。它索引了13亿个视频,每天索引约80万视频。对于ANNS层,平均读取延迟为4秒,平均写入延迟最多为4.84秒。不再需要对完整数据集进行重新训练,无论添加了多少个新数据集,从而消除了节点之间的昂贵数据迁移。迷宫识别具有相似外观和相似音频的重复实时流媒体视频,召回了98%。最重要的是,迷宫也具有成本效益。例如,紧凑的功能设计有助于节省5800 SSD,并且用于运行整个系统的计算资源降低到每十亿视频的250k标准核心。
With the advancement and dominant service of Internet videos, the content-based video deduplication system becomes an essential and dependent infrastructure for Internet video service. However, the explosively growing video data on the Internet challenges the system design and implementation for its scalability in several ways. (1) Although the quantization-based indexing techniques are effective for searching visual features at a large scale, the costly re-training over the complete dataset must be done periodically. (2) The high-dimensional vectors for visual features demand increasingly large SSD space, degrading I/O performance. (3) Videos crawled from the Internet are diverse, and visually similar videos are not necessarily the duplicates, increasing deduplication complexity. (4) Most videos are edited ones. The duplicate contents are more likely discovered as clips inside the videos, demanding processing techniques with close attention to details. To address above-mentioned issues, we propose Maze, a full-fledged video deduplication system. Maze has an ANNS layer that indexes and searches the high dimensional feature vectors. The architecture of the ANNS layer supports efficient reads and writes and eliminates the data migration caused by re-training. Maze adopts the CNN-based feature and the ORB feature as the visual features, which are optimized for the specific video deduplication task. The features are compact and fully reside in the memory. Acoustic features are also incorporated in Maze so that the visually similar videos but having different audio tracks are recognizable. A clip-based matching algorithm is developed to discover duplicate contents at a fine granularity. Maze has been deployed as a production system for two years. It has indexed 1.3 billion videos and is indexing ~800 thousand videos per day. For the ANNS layer, the average read latency is 4 seconds and the average write latency is at most 4.84 seconds. The re-training over the complete dataset is no longer required no matter how many new data sets are added, eliminating the costly data migration between nodes. Maze recognizes the duplicate live streaming videos with both the similar appearance and the similar audio at a recall of 98%. Most importantly, Maze is also cost-effective. For example, the compact feature design helps save 5800 SSDs and the computation resources devoted to running the whole system decrease to 250K standard cores per billion videos.