Evaluation of local spatial-temporal features for cross-view action recognition

Evaluation of local spatial-temporal features for cross-view action recognition
复制标题

跨视图动作识别的局部时空特征评估

DOI:
10.1016/j.neucom.2015.07.105
复制
发表时间:
2016-01
期刊:
影响因子:
6
通讯作者:
Zhang Hua
Zhang Hua
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gao Zan;Nie Weizhi;Liu Anan;Zhang Hua

文献摘要

参考文献

被引文献

相似文献

基于局部时空特征的表示在人类行为识别中非常流行。许多时空显著点检测器和描述符已经被提出。虽然近年来在动作识别方面取得了可喜的成果,但仍存在两个严重的问题:(1)交叉视点动作识别缺乏对局部时空特征的系统评价;(2)对于跨视角动作识别任务,缺乏一种能够自适应桥接多视角不同特征空间的基线方法。本文在可转移字典对学习的框架下,对四种流行的时空特征(STIP、Cuboids、MoSIFT、HoG3D)进行了评价。该框架首先可以在无监督和有监督设置中学习一个可转移的字典对。然后,源视图中的训练样本和目标视图中的测试样本分别用相应的源和目标字典表示,得到稀疏特征表示,用于训练分类器进行动作识别。这样,它可以将不同视图的特征映射到同一特征空间中,以处理跨域任务。在流行的多视图人体动作数据集IXMAS上实现了四个时空特征的评价和可转移字典对学习框架。与代表性方法的对比实验进一步证明了该框架在跨视点人体动作识别上的优越性。
Local spatial–temporal feature-based representation is extremely popular for human action recognition. Many spatial–temporal salient point detectors and descriptors have been proposed. Although the promising results have been achieved for action recognition recently, there still exist two severe problems: (1) there is lack of systematic evaluation of local spatial–temporal features on cross-view action recognition; (2) there is lack of a baseline method especially for the task of cross-view action recognition, which can adaptively bridge different feature spaces from multiple views for this cross-domain task. In this paper, we evaluate four popular spatial–temporal features (STIP, Cuboids, MoSIFT, HoG3D) with the framework of transferable dictionary pair learning. This framework can first learn one transferable dictionary pair in both unsupervised and supervised settings. Then, training samples in the source view and testing samples in the target view can be represented with corresponding source and target dictionary respectively to get sparse feature representations, which is used to training classifier for action recognition. In this way, it can map the features from different views into the same feature space to handle the cross-domain task. The evaluation of four spatial–temporal features and the framework of transferable dictionary pair learning are implemented on the popular multi-view human action dataset, IXMAS. The comparative experiments against the representative methods further demonstrate the superiority of this framework on cross-view human action recognition.
DOI: 10.1007/978-3-540-88682-2_13
发表时间: 2008-10
期刊: --
影响因子: --
作者:
Ali Farhadi;Mostafa Kamali Tabrizi
通讯作者: Ali Farhadi;Mostafa Kamali Tabrizi
DOI: 10.1007/s11263-005-1838-7
发表时间: 2005-09-01
影响因子: 19.5
作者:
Laptev, I
通讯作者: Laptev, I
DOI: 10.1016/j.neucom.2014.04.090
发表时间: 2015-03
期刊: Neurocomputing
影响因子: 6
作者:
Anan Liu;Ning Xu;Yuting Su;Hong Lin;Tong Hao;Zhaoxuan Yang
通讯作者: Anan Liu;Ning Xu;Yuting Su;Hong Lin;Tong Hao;Zhaoxuan Yang
DOI: 10.1109/cvpr.2005.58
发表时间: 2005-06
期刊: 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05)
影响因子: --
作者:
Alper Yilmaz;M. Shah
通讯作者: Alper Yilmaz;M. Shah
DOI: 10.1109/tmm.2012.2237023
发表时间: 2013-04
影响因子: 7.3
作者:
Yi Yang;Zhigang Ma;Alexander Hauptmann;N. Sebe
通讯作者: Yi Yang;Zhigang Ma;Alexander Hauptmann;N. Sebe