Real-time compressed-domain spatiotemporal segmentation and ontologies for video indexing and retrieval

Real-time compressed-domain spatiotemporal segmentation and ontologies for video indexing and retrieval
复制标题

DOI:
10.1109/tcsvt.2004.826768
复制
发表时间:
2004-05
影响因子:
8.4
通讯作者:
V. Mezaris;Y. Kompatsiaris;N. Boulgouris;M. Strintzis
V. Mezaris;Y. Kompatsiaris;N. Boulgouris;M. Strintzis
中科院分区:
工程技术1区
文献类型:
--
作者:
V. Mezaris;Y. Kompatsiaris;N. Boulgouris;M. Strintzis

文献摘要

被引文献

相似文献

本文提出了一种新的图像序列实时压缩域无监督分割算法,并将其应用于视频索引和检索。分割算法使用直接从MPEG-2压缩流中提取的运动和颜色信息。基于双线性运动模型的迭代拒绝方案被用来实现前景/背景分割。在此之后,有意义的前景时空对象的形成,通过最初检查的时间一致性的输出的迭代拒绝,聚类得到的前景宏块连接的区域,最后执行区域跟踪。另外还执行对时空对象的背景分割。MPEG-7兼容的低级别描述符描述的颜色,形状,位置和运动的时空对象被提取,并自动映射到适当的中间级描述符形成一个简单的词汇称为对象本体。这与相关性反馈机制相结合,允许用户查询的高级概念(语义对象,每个由关键字表示)的定性定义和相关视频片段的检索。多关键字查询中的对象之间的期望的空间和时间关系也可以使用镜头本体来表达。实验结果表明,该分割算法的应用已知序列的分割方法的效率。示例查询揭示了采用这种分割算法作为基于对象的视频索引和检索方案的一部分的潜力。
In this paper, a novel algorithm is presented for the real-time, compressed-domain, unsupervised segmentation of image sequences and is applied to video indexing and retrieval. The segmentation algorithm uses motion and color information directly extracted from the MPEG-2 compressed stream. An iterative rejection scheme based on the bilinear motion model is used to effect foreground/background segmentation. Following that, meaningful foreground spatiotemporal objects are formed by initially examining the temporal consistency of the output of iterative rejection, clustering the resulting foreground macroblocks to connected regions and finally performing region tracking. Background segmentation to spatiotemporal objects is additionally performed. MPEG-7 compliant low-level descriptors describing the color, shape, position, and motion of the resulting spatiotemporal objects are extracted and are automatically mapped to appropriate intermediate-level descriptors forming a simple vocabulary termed object ontology. This, combined with a relevance feedback mechanism, allows the qualitative definition of the high-level concepts the user queries for (semantic objects, each represented by a keyword) and the retrieval of relevant video segments. Desired spatial and temporal relationships between the objects in multiple-keyword queries can also be expressed, using the shot ontology. Experimental results of the application of the segmentation algorithm to known sequences demonstrate the efficiency of the proposed segmentation approach. Sample queries reveal the potential of employing this segmentation algorithm as part of an object-based video indexing and retrieval scheme.