Multi-modality video shot clustering with tensor representation

Multi-modality video shot clustering with tensor representation
复制标题

具有张量表示的多模态视频镜头聚类

DOI:
10.1007/s11042-008-0220-5
复制
发表时间:
2008
影响因子:
3.6
通讯作者:
Wu, Fei
Wu, Fei
中科院分区:
计算机科学4区
文献类型:
--
作者:
Liu, Yanan;Wu, Fei

文献摘要

参考文献

被引文献

相似文献

视频分析和理解是当今一个具有挑战性的问题。视频数据具有多种媒体形式,呈现时间序列关联共现(TSAC)特征。传统上,视频被表示为欧几里得空间中的向量。然后,许多学习算法应用于这些向量在高维空间中的降维,分类,聚类和识别以及。然而,视频中的多模态既有各自的特性,又有相互关联的特性,而简单的矢量表示削弱了这些相对独立的模态的能力,甚至在一定程度上忽略了它们之间的关系。聚类是多媒体数据管理的重要技术。最近,一个强大的聚类算法称为亲和传播。在本文中,我们介绍了一个高阶张量框架的视频分析。在这个框架中,我们表示图像帧,音频流和文字记录这三个模态的视频镜头作为数据点的三阶张量。此外,我们提出了一个降维方法的高维特征的视频镜头,明确考虑流形结构的张量空间从时间序列相关的共现多模态媒体数据。我们称之为TensorShot方法。然后,我们利用有效的亲和传播聚类的视频镜头是在张量形式。我们的算法保留了采样张量的子流形的内在结构。在TRECVID 2005新闻视频数据集上的实验结果表明,该算法具有较好的性能。
Video analysis and understanding is a challenging issue nowadays. Video data has multiple media modalities, which present a characteristic of temporal-sequenced associated cooccurrence (TSAC). Traditionally, videos are represented as vectors in the Euclidean space. Many learning algorithms are then applied to these vectors in a high dimensional space for dimensionality reduction, classification, clustering and recognition as well. However, the multiple modalities in video not only have their own properties, but also have correlations between them; whereas the simple vector representation weakens the power of these relatively independent modalities and even ignores their relations to some extent. Clustering is an important technique for multimedia data management. Recently, a powerful clustering algorithm named Affinity Propagation is devised. In this paper, we introduce a higher-order tensor framework for video analysis. In this framework, we represent image frame, audio stream and transcript text which are the three modalities in video shots as data points by the third-order tensor. Besides, we present a dimension reduction method for the high-dimensional features of video shots which explicitly considers the manifold structure of the tensor space from temporal-sequenced associated co-occurring multimodal media data. We call it TensorShot approach. Then we utilize the effective Affinity Propagation to cluster video shots that are in tensor form. Our algorithm preserves the intrinsic structure of the submanifold where tensorshots are sampled. The experiments on TRECVID2005 news video data set show that our algorithm achieves improved performance.
DOI: 10.1109/icassp.2004.1326626
发表时间: 2004-05
期刊: 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子: --
作者:
A. Ekin;Sharath Pankanti;A. Hampapur
通讯作者: A. Ekin;Sharath Pankanti;A. Hampapur
DOI: 10.1145/1101149.1101169
发表时间: 2005-11
期刊: Proceedings of the 13th annual ACM international conference on Multimedia
影响因子: --
作者:
Xiaofei He;Deng Cai;Haifeng Liu;Jiawei Han
通讯作者: Xiaofei He;Deng Cai;Haifeng Liu;Jiawei Han
DOI: 10.1109/icdm.2005.144
发表时间: 2005-11
期刊: Fifth IEEE International Conference on Data Mining (ICDM'05)
影响因子: --
作者:
Ning Liu;Benyu Zhang;Jun Yan;Zheng Chen;Wenyin Liu;F. Bai;Leefeng Chien
通讯作者: Ning Liu;Benyu Zhang;Jun Yan;Zheng Chen;Wenyin Liu;F. Bai;Leefeng Chien
DOI: 10.1109/tmm.2002.802022
发表时间: 2002-12
期刊: IEEE Trans. Multim.
影响因子: --
作者:
C. Ngo;T. Pong;HongJiang Zhang
通讯作者: C. Ngo;T. Pong;HongJiang Zhang
DOI: 10.1109/cvpr.2006.138
发表时间: 2006-06
期刊: 2006 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'06)
影响因子: --
作者:
D. Tao;Xuelong Li;S. Maybank;Xindong Wu
通讯作者: D. Tao;Xuelong Li;S. Maybank;Xindong Wu