Topical video object discovery from key frames by modeling word co-occurrence prior

Topical video object discovery from key frames by modeling word co-occurrence prior
复制标题

DOI:
10.1109/tip.2015.2487834
复制
发表时间:
2015-10
影响因子:
10.6
通讯作者:
Gangqiang Zhao;Junsong Yuan;Gang Hua;Jiong Yang
Gangqiang Zhao;Junsong Yuan;Gang Hua;Jiong Yang
中科院分区:
计算机科学1区
文献类型:
--
作者:
Gangqiang Zhao;Junsong Yuan;Gang Hua;Jiong Yang

文献摘要

被引文献

相似文献

主题视频对象是指视频中经常突出显示的对象。它可以是,例如,产品标识和电视广告中的男主角/女主角。我们提出了一个主题模型,它采用了一个词的共现前有效地发现主题视频对象从一组关键帧。以前的工作使用主题模型,如潜在的Dirichelet分配(LDA),视频对象发现往往需要一袋的视觉词表示,忽略了重要的同现信息之间的本地功能。我们表明,这种数据驱动的共现信息从下而上可以方便地纳入LDA与高斯马尔可夫先验,它结合了自上而下的概率主题建模与自下而上的先验在一个统一的模型。我们对具有挑战性的视频的实验表明,所提出的方法可以发现不同类型的主题对象,尽管在规模,视点,颜色和照明变化,甚至部分遮挡的变化。与没有这种先验知识的主题模型相比,同现先验知识的有效性得到了清楚的证明。
A topical video object refers to an object, that is, frequently highlighted in a video. It could be, e.g., the product logo and the leading actor/actress in a TV commercial. We propose a topic model that incorporates a word co-occurrence prior for efficient discovery of topical video objects from a set of key frames. Previous work using topic models, such as latent Dirichelet allocation (LDA), for video object discovery often takes a bag-of-visual-words representation, which ignored important co-occurrence information among the local features. We show that such data driven co-occurrence information from bottom-up can conveniently be incorporated in LDA with a Gaussian Markov prior, which combines top-down probabilistic topic modeling with bottom-up priors in a unified model. Our experiments on challenging videos demonstrate that the proposed approach can discover different types of topical objects despite variations in scale, view-point, color and lighting changes, or even partial occlusions. The efficacy of the co-occurrence prior is clearly demonstrated when compared with topic models without such priors.