Set-Constrained Viterbi for Set-Supervised Action Segmentation

Set-Constrained Viterbi for Set-Supervised Action Segmentation
复制标题

DOI:
10.1109/cvpr42600.2020.01083
复制
发表时间:
2020-02
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Jun Li;S. Todorovic
Jun Li;S. Todorovic
中科院分区:
其他
文献类型:
--
作者:
Jun Li;S. Todorovic

文献摘要

被引文献

相似文献

这篇论文是关于弱监督动作分割的,其中基础真值只指定了训练视频中出现的一组动作,而不是它们的真实时间顺序。先前的工作通常使用独立标记视频帧的分类器来生成伪基础真理,并使用多实例学习来训练分类器。我们通过指定HMM来扩展这个框架,HMM解释了动作类的共同出现和它们的时间长度,并通过基于viterbi的损失显式地训练HMM。我们的第一个贡献是提出了一个新的集约束Viterbi算法(SCV)。给定一个视频,SCV生成MAP动作分割,满足基本事实。在我们的HMM训练中,这个预测被用作一种基于帧的伪基础真值。我们在训练方面的第二个贡献是对共享相同动作类的训练视频之间的特征亲和力进行了新的正则化。在Breakfast, MPII Cooking2, Hollywood Extended数据集上对动作分割和对齐的评估表明,我们在这两个任务上的性能比之前的工作有了显著的提高。
This paper is about weakly supervised action segmentation, where the ground truth specifies only a set of actions present in a training video, but not their true temporal ordering. Prior work typically uses a classifier that independently labels video frames for generating the pseudo ground truth, and multiple instance learning for training the classifier. We extend this framework by specifying an HMM, which accounts for co-occurrences of action classes and their temporal lengths, and by explicitly training the HMM on a Viterbi-based loss. Our first contribution is the formulation of a new set-constrained Viterbi algorithm (SCV). Given a video, the SCV generates the MAP action segmentation that satisfies the ground truth. This prediction is used as a framewise pseudo ground truth in our HMM training. Our second contribution in training is a new regularization of feature affinities between training videos that share the same action classes. Evaluation on action segmentation and alignment on the Breakfast, MPII Cooking2, Hollywood Extended datasets demonstrates our significant performance improvement for the two tasks over prior work.