Group Sparse-Based Mid-Level Representation for Action Recognition

Group Sparse-Based Mid-Level Representation for Action Recognition
复制标题

DOI:
10.1109/tsmc.2016.2625840
复制
发表时间:
2017-04
期刊:
IEEE Transactions on Systems, Man, and Cybernetics: Systems
影响因子:
--
通讯作者:
Shiwei Zhang;Changxin Gao;Feifei Chen;Sihui Luo;N. Sang
Shiwei Zhang;Changxin Gao;Feifei Chen;Sihui Luo;N. Sang
中科院分区:
其他
文献类型:
--
作者:
Shiwei Zhang;Changxin Gao;Feifei Chen;Sihui Luo;N. Sang

文献摘要

被引文献

相似文献

中层部分被证明对视频中的人类动作识别是有效的。通常,这些语义部分首先用一些启发式规则进行挖掘,然后用体积最大池(VMP)方法表示视频。然而,这些方法有两个问题:1)VMP策略按静态网格划分视频。在这种情况下,语义部分可能在不同的视频中出现在不同的本地化中。这意味着VMP策略失去了时空不变性。为了解决这个问题,我们提出了一种显著驱动的最大池方案来表示视频。我们通过显著图提取视频语义线索,并动态汇集局部最大响应。该方案可以被认为是一种基于语义内容的特征对齐方法,2)启发式规则发现的部分可能是直观的,但不足以区分动作分类,因为它们忽略了检测器之间的关系。针对这个问题,我们提出了一种稀疏分类器模型来选择有区别的部分。此外,为了进一步提高表示的区分能力,我们提出了根据模型系数对应的输入大小进行特征选择。我们在四个具有挑战性的数据集上进行了实验--KTH、奥林匹克体育、UCF50和HMDB51。实验结果表明,该方法的性能明显优于现有的方法。
Mid-level parts are shown to be effective for human action recognition in videos. Typically, these semantic parts are first mined with some heuristic rules, then videos are represented via volumetric max-pooling (VMP) method. However, these methods have two issues: 1) the VMP strategy divides videos by static grids. In this case, a semantic part may occur in different localizations in different videos. That means the VMP strategy loses the space-time invariance. To solve this problem, we propose to apply a saliency-driven max-pooling scheme to represent a video. We extract the video semantic cues by the saliency map, and dynamically pool the local maximum responses. This scheme can be considered as a semantic content-based feature alignment method and 2) the parts discovered by heuristic rules may be intuitive but not discriminative enough for action classification because they neglect the relations between the detectors. For this issue, we propose to apply a sparse classifier model to select discriminative parts. Moreover, to further improve the discriminative ability of the representation, we propose to conduct feature selection by the corresponding entry magnitude of the model coefficients. We conduct experiments on four challenging datasets—KTH, Olympic Sports, UCF50, and HMDB51. The results show that the proposed method significantly outperforms the state-of-the-art methods.