Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos

Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos
复制标题

传输潜在支持向量机用于视频中动作的联合识别和定位

DOI:
10.1109/tcyb.2015.2482970
复制
发表时间:
2016
影响因子:
11.8
通讯作者:
Jia Yunde
Jia Yunde
中科院分区:
计算机科学1区
文献类型:
--
作者:
Liu Cuiwei;Wu Xinxiao;Jia Yunde

文献摘要

相似文献

在本文中,我们开发了一种新的转移潜在的支持向量机的联合识别和定位的行动,使用Web图像和弱注释的训练视频。该模型将仅用动作标签标注的训练视频作为输入,以减轻手动标注动作位置的费力和耗时。由于视频中动作位置的地面实况不可用,因此在我们的方法中,位置被建模为潜在变量,并在训练和测试阶段进行推断。为了提高定位精度与动作位置的一些先验信息的目的,我们收集了一些Web图像,这是注释与动作标签和动作位置学习判别模型,通过加强视频和Web图像之间的局部相似性。为了处理Web图像和视频的异构特征,采用基于随机聚类森林的结构化变换将Web图像映射到视频。在两个公开的动作数据集上的实验证明了该模型对动作定位和动作识别的有效性。
In this paper, we develop a novel transfer latent support vector machine for joint recognition and localization of actions by using Web images and weakly annotated training videos. The model takes training videos which are only annotated with action labels as input for alleviating the laborious and time-consuming manual annotations of action locations. Since the ground-truth of action locations in videos are not available, the locations are modeled as latent variables in our method and are inferred during both training and testing phrases. For the purpose of improving the localization accuracy with some prior information of action locations, we collect a number of Web images which are annotated with both action labels and action locations to learn a discriminative model by enforcing the local similarities between videos and Web images. A structural transformation based on randomized clustering forest is used to map the Web images to videos for handling the heterogeneous features of Web images and videos. Experiments on two public action datasets demonstrate the effectiveness of the proposed model for both action localization and action recognition.