Video linkage: group based copied video detection

Video linkage: group based copied video detection
复制标题

DOI:
10.1145/1386352.1386404
复制
发表时间:
2008-07
期刊:
--
影响因子:
--
通讯作者:
Hung-sik Kim;JeongKyu Lee;Haibin Liu;Dongwon Lee
Hung-sik Kim;JeongKyu Lee;Haibin Liu;Dongwon Lee
中科院分区:
其他
文献类型:
--
作者:
Hung-sik Kim;JeongKyu Lee;Haibin Liu;Dongwon Lee

文献摘要

被引文献

相似文献

近年来,可以共享由用户创建的视频片段(例如YouTube和Yahoo视频)的网站。但是,此类网站的挑战之一是防止通过非法复制和编辑其他视频的场景来侵犯版权的视频片段。由于每天上传的剪辑数量庞大,因此需要自动检测(非法)复制的视频片段。对于这个问题,在本文中,我们提出了一个新颖的框架,称为视频链接,该框架基于记录链接技术。我们的建议基于以下观察结果:(1)视频剪辑可以表示为关键帧的“组”,(2)两个视频剪辑被认为是相似的,如果两组关键帧与整体相似 - 即,可以通过基于图的相似性度量(例如最大基数双方匹配)来测量两个视频剪辑的相似性,并且(3)如果将视频剪辑VA复制到VB,那么VA和VB必须以某种方式相似,但并非所有类似的视频剪辑都是非法复制的 - 即,类似的视频可以用作快速检测复制视频的过滤器。使用真实和合成数据集(即,平均而言,我们的提案达到0.94的精度和0.93的召回方式,我们的观察结果和视频链接技术的有效性得到了彻底评估,即我们的提案达到0.94,并以0.93的召回方式达到0.93。
Sites to share user-created video clips such as YouTube and Yahoo Video have become greatly popular in recent years. One of the challenges of such sites is, however, to prevent video clips that violate copyrights by illegally copying and editing scenes from other videos. Due to the sheer number of clips uploaded every day, automatic methods to detect (illegally) copied video clips in a large collection are desirable. Toward this problem, in this paper, we present a novel framework, termed as Video Linkage, that is based on the record linkage techniques. Our proposal is based on the observations that: (1) a video clip can be represented as a "group" of key frames, (2) two video clips are deemed to be similar if two groups of key frames are similar as a whole - i.e., the similarity of two video clips can be measured by means of graph-based similarity measures such as maximal cardinality bipartite matching, and (3) if a video clip va is copied to vb, then va and vb must be somehow similar, but not all similar video clips are illegally copied ones - i.e., similar videos can be used as a filter for fast detection of copied videos. The validity of our observations and Video Linkage technique is thoroughly evaluated using both real and synthetic data sets - i.e., on average, our proposals achieved 0.94 as precision and 0.93 as recall across 10 genres and 6 editing patterns.