Unsupervised learning of human action categories using spatial-temporal words

Unsupervised learning of human action categories using spatial-temporal words
复制标题

DOI:
10.1007/s11263-007-0122-4
复制
发表时间:
2008-09-01
影响因子:
19.5
通讯作者:
Fei-Fei, Li
Fei-Fei, Li
中科院分区:
计算机科学2区
文献类型:
--
作者:
Niebles, Juan Carlos;Wang, Hongcheng;Fei-Fei, Li

文献摘要

被引文献

相似文献

我们为人类行动类别提供了一种新颖的无监督学习方法。视频序列通过提取时空兴趣点表示为空间词的集合。该算法会自动学习时空单词的概率分布和与人类作用类别相对应的中间主题。这是通过使用潜在主题模型(例如概率潜在语义分析(PLSA)模型和潜在的Dirichlet分配(LDA)来实现的。由于应用概率模型的应用,我们的方法可以处理由动态背景和移动相机产生的嘈杂特征点。给定一个新颖的视频序列,该算法可以对视频中包含的人类作用进行分类和定位。我们在三个具有挑战性的数据集上测试了算法:KTH人类运动数据集,Weizmann人类动作数据集和最新的花样滑冰动作数据集。我们的结果反映了这种简单方法的希望。此外,我们的算法可以识别并将多个动作定位在包含多个动作的长而复杂的视频序列中。
We present a novel unsupervised learning method for human action categories. A video sequence is represented as a collection of spatial-temporal words by extracting space-time interest points. The algorithm automatically learns the probability distributions of the spatial-temporal words and the intermediate topics corresponding to human action categories. This is achieved by using latent topic models such as the probabilistic Latent Semantic Analysis (pLSA) model and Latent Dirichlet Allocation (LDA). Our approach can handle noisy feature points arisen from dynamic background and moving cameras due to the application of the probabilistic models. Given a novel video sequence, the algorithm can categorize and localize the human action(s) contained in the video. We test our algorithm on three challenging datasets: the KTH human motion dataset, the Weizmann human action dataset, and a recent dataset of figure skating actions. Our results reflect the promise of such a simple approach. In addition, our algorithm can recognize and localize multiple actions in long and complex video sequences containing multiple motions.