n-gram Models for Video Semantic Indexing

n-gram Models for Video Semantic Indexing
复制标题

DOI:
10.1145/2647868.2654961
复制
发表时间:
2014-11
期刊:
Proceedings of the 22nd ACM international conference on Multimedia
影响因子:
--
通讯作者:
Nakamasa Inoue;K. Shinoda
Nakamasa Inoue;K. Shinoda
中科院分区:
其他
文献类型:
--
作者:
Nakamasa Inoue;K. Shinoda

文献摘要

相似文献

我们提出了用于视频语义索引的镜头序列的n-gram建模,其中从视频镜头中提取语义概念。大多数以前的研究都假设视频剪辑中的视频镜头是相互独立的。假设n个连续的视频镜头是相关的,我们对它们之间的时间依赖性进行建模。我们的模型通过有效地使用来自先前视频镜头的信息来提高对遮挡和摄像机角度变化的鲁棒性。在我们对TRECVID 2012语义索引基准的实验中,我们将所提出的模型应用于使用高斯混合模型和支持向量机的系统。平均精度从30.62%提高到32.14%,这是我们所知的TRECVID 2012语义索引的最佳性能。
We propose n-gram modeling of shot sequences for video semantic indexing, in which semantic concepts are extracted from a video shot. Most previous studies for this task have assumed that video shots in a video clip are independent from each other. We model the time-dependency between them assuming that n-consecutive video shots are dependent. Our models improve the robustness against occlusion and camera-angle changes by effectively using information from the previous video shots. In our experiments on the TRECVID 2012 Semantic Indexing Benchmark, we applied the proposed models to a system using Gaussian mixture models and support vector machines. Mean average precision was improved from 30.62% to 32.14%, which is the best performance on the TRECVID 2012 Semantic Indexing to the best of our knowledge.