Recognizing 50 human action categories of web videos

Recognizing 50 human action categories of web videos
复制标题

DOI:
10.1007/s00138-012-0450-4
复制
发表时间:
2013-07-01
影响因子:
3.3
通讯作者:
Shah, Mubarak
Shah, Mubarak
中科院分区:
计算机科学4区
文献类型:
--
作者:
Reddy, Kishore K.;Shah, Mubarak

文献摘要

被引文献

相似文献

与KTH(6个动作),IXMAS(13个动作)和Weizmann(10个动作)等数据集相比,从网络上拍摄的大类无约束视频的动作识别是一个非常具有挑战性的问题。像摄像机运动、不同的视角、大的类间变化、杂乱的背景、遮挡、不良的照明条件和网络视频质量差等挑战导致大多数最先进的动作识别方法失败。此外,类别数量的增加和列入的行动非常混乱,增加了挑战。在本文中,我们建议使用场景上下文信息从移动和静止的像素在关键帧中,结合运动特征,以解决动作识别问题的一个大的(50个动作)数据集与视频从网络上。我们对多个特征执行早期和晚期融合的组合,以处理非常大量的类别。我们证明了场景上下文是在非常大的数据集上执行动作识别的一个非常重要的特征。所提出的方法不需要任何类型的视频稳定,人检测,或跟踪和修剪的功能。我们的方法在大量的动作类别上具有良好的性能;它已经在包含50个动作类别的UCF 50数据集上进行了测试,该数据集是包含11个动作类别的UCF YouTube Action(UCF 11)数据集的扩展。我们还在KTH和HMDB 51数据集上测试了我们的方法以进行比较。
Action recognition on large categories of unconstrained videos taken from the web is a very challenging problem compared to datasets like KTH (6 actions), IXMAS (13 actions), and Weizmann (10 actions). Challenges like camera motion, different viewpoints, large interclass variations, cluttered background, occlusions, bad illumination conditions, and poor quality of web videos cause the majority of the state-of-the-art action recognition approaches to fail. Also, an increased number of categories and the inclusion of actions with high confusion add to the challenges. In this paper, we propose using the scene context information obtained from moving and stationary pixels in the key frames, in conjunction with motion features, to solve the action recognition problem on a large (50 actions) dataset with videos from the web. We perform a combination of early and late fusion on multiple features to handle the very large number of categories. We demonstrate that scene context is a very important feature to perform action recognition on very large datasets. The proposed method does not require any kind of video stabilization, person detection, or tracking and pruning of features. Our approach gives good performance on a large number of action categories; it has been tested on the UCF50 dataset with 50 action categories, which is an extension of the UCF YouTube Action (UCF11) dataset containing 11 action categories. We also tested our approach on the KTH and HMDB51 datasets for comparison.