A Weakly Supervised Multi-task Ranking Framework for Actor–Action Semantic Segmentation

A Weakly Supervised Multi-task Ranking Framework for Actor–Action Semantic Segmentation
复制标题

DOI:
10.1007/s11263-019-01244-7
复制
发表时间:
2019-10
影响因子:
19.5
通讯作者:
Yan Yan-Yan;Chenliang Xu;Dawen Cai;Jason J. Corso
Yan Yan-Yan;Chenliang Xu;Dawen Cai;Jason J. Corso
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yan Yan-Yan;Chenliang Xu;Dawen Cai;Jason J. Corso

文献摘要

被引文献

相似文献

人类行为和活动模式的建模近年来引起了人们极大的研究兴趣。为了准确地模拟人类行为,我们需要在视频中执行细粒度的人类活动理解。视频中的细粒度活动理解最近引起了相当大的关注,从动作分类转向详细的演员和动作理解,为尖端自主系统的感知需求提供了令人信服的结果。然而,目前用于详细了解行动者和动作的方法有很大的局限性:它们需要大量精细标记的数据,并且它们无法捕获行动者和动作之间的任何内部关系。为了解决这些问题,在本文中,我们提出了一种新的Schattenp-norm鲁棒多任务排序模型,用于弱监督的演员-动作分割,其中仅为训练样本提供视频级标签。我们的模型能够在不同的演员和动作之间共享有用的信息,同时学习一个排序矩阵来分别为演员和动作选择有代表性的超体素。最后的分割结果由一个条件随机场生成,该随机场考虑了视频部分的各种排名分数。在演员-动作数据集和youtube -对象数据集上的大量实验结果表明,所提出的方法优于最先进的弱监督方法,并且与性能最好的完全监督方法一样好。
Modeling human behaviors and activity patterns has attracted significant research interest in recent years. In order to accurately model human behaviors, we need to perform fine-grained human activity understanding in videos. Fine-grained activity understanding in videos has attracted considerable recent attention with a shift from action classification to detailed actor and action understanding that provides compelling results for perceptual needs of cutting-edge autonomous systems. However, current methods for detailed understanding of actor and action have significant limitations: they require large amounts of finely labeled data, and they fail to capture any internal relationship among actors and actions. To address these issues, in this paper, we propose a novel Schattenp-norm robust multi-task ranking model for weakly-supervised actor–action segmentation where only video-level tags are given for training samples. Our model is able to share useful information among different actors and actions while learning a ranking matrix to select representative supervoxels for actors and actions respectively. Final segmentation results are generated by a conditional random field that considers various ranking scores for video parts. Extensive experimental results on both the actor–action dataset and the Youtube-objects dataset demonstrate that the proposed approach outperforms the state-of-the-art weakly supervised methods and performs as well as the top-performing fully supervised method.