Unsupervised Action Discovery and Localization in Videos

Unsupervised Action Discovery and Localization in Videos
复制标题

DOI:
10.1109/iccv.2017.82
复制
发表时间:
2017-10
期刊:
2017 IEEE International Conference on Computer Vision (ICCV)
影响因子:
--
通讯作者:
K. Soomro;M. Shah
K. Soomro;M. Shah
中科院分区:
其他
文献类型:
--
作者:
K. Soomro;M. Shah

文献摘要

被引文献

相似文献

本文是第一个解决视频中的无监督动作定位问题。给定没有边界框注释的未标记数据,我们提出了一种新的方法:1)发现动作类标签和2)时空定位视频中的动作。它首先计算本地视频特征,以在一组未标记的训练视频上应用谱聚类。对于每个视频簇,构建一个无向图来提取一个主导集,这是已知的高内部同质性和外部顶点之间的不同质性。接下来,应用判别聚类方法,通过训练每个簇的分类器,迭代地从非主导集中选择视频,并获得完整的视频动作类。一旦发现类别,则通过首先将每个发现的类别中的视频过分割成超体素并构建有向图来应用具有时间约束的背包问题的变体,来选择每个集群内的训练视频以执行自动时空注释。背包优化联合收集超体素的子集,通过强制注释的动作是时空连接的,其体积是演员的大小。这些注释用于训练SVM动作分类器。在测试过程中,使用类似的背包方法来定位动作,其中将超体素分组在一起,并使用从发现的动作类中学习的视频来识别这些动作。我们在UCF-Sports,Sub-JHMDB,JHMDB,THUMOS 13和UCF 101数据集上评估了我们的方法。我们的实验表明,尽管没有使用动作类标签和边界框注释,我们仍然能够获得与最先进的监督方法相竞争的结果。
This paper is the first to address the problem of unsupervised action localization in videos. Given unlabeled data without bounding box annotations, we propose a novel approach that: 1) Discovers action class labels and 2) Spatio-temporally localizes actions in videos. It begins by computing local video features to apply spectral clustering on a set of unlabeled training videos. For each cluster of videos, an undirected graph is constructed to extract a dominant set, which are known for high internal homogeneity and in-homogeneity between vertices outside it. Next, a discriminative clustering approach is applied, by training a classifier for each cluster, to iteratively select videos from the non-dominant set and obtain complete video action classes. Once classes are discovered, training videos within each cluster are selected to perform automatic spatio-temporal annotations, by first over-segmenting videos in each discovered class into supervoxels and constructing a directed graph to apply a variant of knapsack problem with temporal constraints. Knapsack optimization jointly collects a subset of supervoxels, by enforcing the annotated action to be spatio-temporally connected and its volume to be the size of an actor. These annotations are used to train SVM action classifiers. During testing, actions are localized using a similar Knapsack approach, where supervoxels are grouped together and SVM, learned using videos from discovered action classes, is used to recognize these actions. We evaluate our approach on UCF-Sports, Sub-JHMDB, JHMDB, THUMOS13 and UCF101 datasets. Our experiments suggest that despite using no action class labels and no bounding box annotations, we are able to get competitive results to the state-of-the-art supervised methods.