Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos
Transfer Latent SVM for Joint Recognition and Localization of Actions in Videos
复制标题
传输潜在支持向量机用于视频中动作的联合识别和定位
DOI:
10.1109/tcyb.2015.2482970
复制
发表时间:
2016
影响因子:
11.8
通讯作者:
Jia Yunde
中科院分区:
文献类型:
--
作者:
Liu Cuiwei;Wu Xinxiao;Jia Yunde
In this paper, we develop a novel transfer latent support vector machine for joint recognition and localization of actions by using Web images and weakly annotated training videos. The model takes training videos which are only annotated with action labels as input for alleviating the laborious and time-consuming manual annotations of action locations. Since the ground-truth of action locations in videos are not available, the locations are modeled as latent variables in our method and are inferred during both training and testing phrases. For the purpose of improving the localization accuracy with some prior information of action locations, we collect a number of Web images which are annotated with both action labels and action locations to learn a discriminative model by enforcing the local similarities between videos and Web images. A structural transformation based on randomized clustering forest is used to map the Web images to videos for handling the heterogeneous features of Web images and videos. Experiments on two public action datasets demonstrate the effectiveness of the proposed model for both action localization and action recognition.