Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task Videos

Semi-Weakly-Supervised Learning of Complex Actions from Instructional Task Videos
复制标题

DOI:
10.1109/cvpr52688.2022.00334
复制
发表时间:
2022-06
期刊:
2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yuhan Shen;Ehsan Elhamifar
Yuhan Shen;Ehsan Elhamifar
中科院分区:
其他
文献类型:
--
作者:
Yuhan Shen;Ehsan Elhamifar

文献摘要

相似文献

针对教学任务视频中的动作分割问题,我们使用少量的弱标记训练视频和大量的未标记视频,我们称之为动作的半弱监督学习(SWSL)。我们提出了一个通用的SWSL框架,它可以有效地学习这两种类型的视频,并可以利用现有的任何弱监督动作分割方法。我们的主要观察是,未标记的视频和同一任务中标记较弱的视频之间的距离很小,但通常不是零。因此,我们开发了一种软限制编辑(SRE)损失,以鼓励同一任务的未标记视频的预测转录和弱标记视频的地面真实转录之间的微小差异。为了计算SRE损失,我们开发了一种灵活的转录预测(FTP)方法,该方法使用动作分类器的输出来找出在未标记的视频中发生的转录的长度和动作的序列。我们提出了一种有效的学习方案,我们在最小化建议的损失和为未标记的视频生成伪转录之间交替进行。通过在两个基准数据集上的实验,我们证明了我们的方法可以显著提高使用未标记视频的性能,特别是在弱标记视频数量较少的情况下。11代码可在https://github.com/Yuhan-Shen/SWSL..上找到
We address the problem of action segmentation in instructional task videos with a small number of weakly-labeled training videos and a large number of unlabeled videos, which we refer to as Semi-Weakly-Supervised Learning (SWSL) of actions. We propose a general SWSL framework that can efficiently learn from both types of videos and can leverage any of the existing weakly-supervised action segmentation methods. Our key observation is that the distance between the transcript of an unlabeled video and those of the weakly-labeled videos from the same task is small yet often nonzero. Therefore, we develop a Soft Restricted Edit (SRE) loss to encourage small variations between the predicted transcripts of unlabeled videos and ground-truth transcripts of the weakly-labeled videos of the same task. To compute the SRE loss, we develop a flexible transcript prediction (FTP) method that uses the output of the action classifier to find both the length of the transcript and the sequence of actions occurring in an unlabeled video. We propose an efficient learning scheme in which we alternate between minimizing our proposed loss and generating pseudo-transcripts for unlabeled videos. By experiments on two benchmark datasets, we demonstrate that our approach can significantly improve the performance by using unlabeled videos, especially when the number of weakly-labeled videos is small. 11Code available at https://github.com/Yuhan-Shen/SWSL..