Two-Stream SR-CNNs for Action Recognition in Videos

Two-Stream SR-CNNs for Action Recognition in Videos
复制标题

DOI:
10.5244/c.30.108
复制
发表时间:
2016
影响因子:
5.4
通讯作者:
Yifan Wang;Jie Song;Limin Wang;L. Gool;Otmar Hilliges
Yifan Wang;Jie Song;Limin Wang;L. Gool;Otmar Hilliges
中科院分区:
材料科学3区
文献类型:
--
作者:
Yifan Wang;Jie Song;Limin Wang;L. Gool;Otmar Hilliges

文献摘要

被引文献

相似文献

人类动作是计算机视觉研究和理解中的一个高级概念,它可能贝内于不同的语义,如人类姿势,交互对象和场景上下文。在本文中,我们明确地利用语义线索与现有的人/物体检测器的动作识别视频,并深入研究其对不同类型的动作识别性能的影响。具体来说,我们提出了一种新的深度架构,将人/物体检测结果纳入框架,称为基于双流语义区域的CNN(SR-CNN)。我们提出的架构不仅与原始的双流CNN共享强大的建模能力,而且还表现出利用语义线索(例如场景,人,对象)进行动作理解的灵活性。我们在UCF 101数据集上进行了实验,并证明了其上级性能优于原始的双流CNN。此外,我们还系统地研究了语义线索对不同类型动作类识别性能的影响,并试图为建立更合理的动作基准和开发更好的识别算法提供一些见解。
Human action is a high-level concept in computer vision research and understanding it may benefit from different semantics, such as human pose, interacting objects, and scene context. In this paper, we explicitly exploit semantic cues with aid of existing human/object detectors for action recognition in videos, and thoroughly study their effect on the recognition performance for different types of actions. Specifically, we propose a new deep architecture by incorporating human/object detection results into the framework, called two-stream semantic region based CNNs (SR-CNNs). Our proposed architecture not only shares great modeling capacity with the original two-stream CNNs, but also exhibits the flexibility of leveraging semantic cues (e.g. scene, person, object) for action understanding. We perform experiments on the UCF101 dataset and demonstrate its superior performance to the original two-stream CNNs. In addition, we systematically study the effect of incorporating semantic cues on the recognition performance for different types of action classes, and try to provide some insights for building more reasonable action benchmarks and developing better recognition algorithms.