ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video

ML-HDP: A Hierarchical Bayesian Nonparametric Model for Recognizing Human Actions in Video
复制标题

DOI:
10.1109/tcsvt.2018.2816960
复制
发表时间:
2019-03-01
影响因子:
8.4
通讯作者:
Lee, Young-Koo
Lee, Young-Koo
中科院分区:
工程技术1区
文献类型:
--
作者:
Nguyen Anh Tu;Thien Huynh-The;Lee, Young-Koo

文献摘要

被引文献

相似文献

视频中的动作识别是计算机视觉研究的一个重要领域,因为它具有各种应用,从视觉监控到人机交互。为了解决动作识别问题,本文提出了一个框架,联合建模多个复杂的动作和运动单元在不同的层次结构。我们提出了一个生成式的主题模型,即多标签层次狄利克雷过程(ML-HDP)。ML-HDP模型制定了动作和运动单元的同现关系,并实现了高精度的识别。特别是,我们的主题模型具有三级表示的动作理解,其中低级别的本地功能连接到高级别的动作通过中级原子动作。这允许识别模型有区别地工作。在我们的ML-HDP中,原子动作被视为潜在的主题,并自动从数据中发现。此外,我们以半监督的方式将类别标签的概念纳入我们的模型中,以有效地学习和推断多标签视频。使用发现的主题和推断的标签,这是联合分配给本地功能,我们提出了直接的方法来执行三个识别任务,包括动作分类,联合分类和分割的连续动作,时空动作本地化。在实验中,我们探索了三种不同特征的使用,并在四个公共数据集上证明了我们提出的方法对这些任务的有效性:KTH,MSR-II,Hollywood 2和UCF 101。
Action recognition from videos is an important area of computer vision research due to its various applications, ranging from visual surveillance to human-computer interaction. To address action recognition problems, this paper presents a framework that jointly models multiple complex actions and motion units at different hierarchical levels. We achieve this by proposing a generative topic model, namely, multi-label hierarchical Dirichlet process (ML-HDP). The ML-HDP model formulates the co-occurrence relationship of actions and motion units, and enables highly accurate recognition. In particular, our topic model possesses the three-level representation in action understanding, where low-level local features are connected to high-level actions via mid-level atomic actions. This allows the recognition model to work discriminatively. In our ML-HDP, atomic actions are treated as latent topics and automatically discovered from data. In addition, we incorporate the notion of class labels into our model in a semi-supervised fashion to effectively learn and infer multi-labeled videos. Using discovered topics and inferred labels, which are jointly assigned to local features, we present the straightforward methods to perform three recognition tasks including action classification, joint classification and segmentation of continuous actions, and spatiotemporal action localization. In experiments, we explore the use of three different features and demonstrate the effectiveness of our proposed approach for these tasks on four public datasets: KTH, MSR-II, Hollywood2, and UCF101.