Video content categorization using the double decomposition

Video content categorization using the double decomposition
复制标题

DOI:
10.1007/s11042-012-1213-y
复制
发表时间:
2013-10
影响因子:
3.6
通讯作者:
Youtian Du;Feng Chen;Wenli Xu;Xueming Qian
Youtian Du;Feng Chen;Wenli Xu;Xueming Qian
中科院分区:
计算机科学4区
文献类型:
--
作者:
Youtian Du;Feng Chen;Wenli Xu;Xueming Qian

文献摘要

相似文献

由于所涉及的组件和事件的多样性,视频内容包含复杂的结构。例如,监控视频通常记录多对象交互,并包含各种尺度的运动细节;网络视频是由多模态线索组成的,每个线索通常由不同尺度的信息组成。视频内容通常包含两种内在结构的组合:多模态/多尺度和多对象/多尺度。因此,本文提出了一种新的视频内容建模框架,在该框架下,针对每种类型的结构组合,通过双重分解将视频内容分解为多个相互作用的过程。为了对结果过程进行建模,我们提出了一种称为双分解隐马尔可夫模型(ddhmm)的方法。ddhmm包含多个状态链,这些状态链对应于交互过程。为了使每条链的状态切换频率与相应过程的尺度一致,在ddhmm中引入了持续状态变量。该方法可以很好地模拟各过程之间的相互作用关系和各过程之间的动力学关系。我们在提出的框架下讨论了适当的特征,并在人体运动识别和网络视频分类两种应用中评估了ddhmm。实验结果表明,在两种情况下,双重分解都提高了视频分类的性能。
Video contents contain complex structures due to the variety of the components and events involved. For example, surveillance videos often record multi-object interactions and consist of various scales of motion detail; Web videos are composed of multimodal cues, and each cue generally consists of a variety of scales of information. Generally, video contents comprise two types of the combination of the inherent structures: multi-modality/multi-scale and multi-object /multi-scale. Therefore, in this paper, we propose a new framework for video content modeling, under which video contents are decomposed into multiple interacting processes by double decomposition that aims at each type of combination of structures. To model the resulting processes, we propose a method named double-decomposed hidden Markov models (DDHMMs). DDHMMs contain multiple state chains that correspond to the interacting processes. To make the switching frequency of states in each chain consistent with the scale of the corresponding process, a durational state variable is introduced in DDHMMs. The proposed method performs well in modeling the relations among the interacting processes and the dynamics of each. We discuss the appropriate features under the proposed framework and evaluate DDHMMs in two applications, human motion recognition and web video categorization. The experimental results demonstrate that the double decomposition enhances video categorization performance in both cases.