Enabling Weakly Supervised Temporal Action Localization From On-Device Learning of the Video Stream

Enabling Weakly Supervised Temporal Action Localization From On-Device Learning of the Video Stream
复制标题

DOI:
10.1109/tcad.2022.3197536
复制
发表时间:
2022-08
影响因子:
2.9
通讯作者:
Yue Tang;Yawen Wu;Peipei Zhou;Jingtong Hu
Yue Tang;Yawen Wu;Peipei Zhou;Jingtong Hu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yue Tang;Yawen Wu;Peipei Zhou;Jingtong Hu

文献摘要

相似文献

视频中的动作检测已广泛应用于设备端应用,如汽车、机器人等。实际的设备端视频总是包含动作和背景且未经剪辑。理想的模型应既能识别动作类别,又能定位动作发生的时间位置。这样的任务被称为时间动作定位(TAL),它通常在云端进行训练,在云端会收集并标注多个未经剪辑的视频。理想的TAL模型应能持续且本地化地从新数据中学习,这既能直接提高动作检测精度,又能保护用户隐私。然而,直接在设备上训练TAL模型并非易事。要训练一个能精确识别和定位每个动作的TAL模型,需要大量带有时间标注的视频样本。然而,逐帧标注视频极其耗时且昂贵。尽管弱监督时间动作定位(W - TAL)已被提出用于从仅有视频级标签的未经剪辑视频中学习,但这种方法也不适用于设备端学习场景。在实际的设备端学习应用中,数据是以流的形式收集的。例如,设备上的摄像头会持续数小时或数天收集视频帧,几乎所有类别的动作都包含在一个长长的视频流中。将这样一个长视频流分割成多个视频片段需要大量人力,这阻碍了将TAL任务应用于实际设备端学习应用的探索。为了使W - TAL模型能从一个长的、未经剪辑的视频流中学习,我们提出了一种高效的视频学习方法,它能直接适应新环境。我们首先提出一种自适应视频分割方法以及一种基于对比分数的片段合并方法,将视频流转换为多个片段。然后,我们探索TAL任务的不同采样策略,以尽可能少地请求标签。据我们所知,我们是首次尝试直接从设备端的长视频流中学习。在THUMOS’14数据集上的实验结果表明,我们的方法在无需任何费力的手动视频分割的情况下,性能与当前W - TAL的最先进水平(SOTA)相当。
Detecting actions in videos have been widely applied in on-device applications, such as cars, robots, etc. Practical on-device videos are always untrimmed with both action and background. It is desirable for a model to both recognize the class of action and localize the temporal position where the action happens. Such a task is called temporal action location (TAL), which is always trained on the cloud where multiple untrimmed videos are collected and labeled. It is desirable for a TAL model to continuously and locally learn from new data, which can directly improve the action detection precision while protecting customers’ privacy. However, directly training a TAL model on the device is nontrivial. To train a TAL model which can precisely recognize and localize each action, tremendous video samples with temporal annotations are required. However, annotating videos frame by frame is exorbitantly time consuming and expensive. Although weakly supervised temporal action localization (W-TAL) has been proposed to learn from untrimmed videos with only video-level labels, such an approach is also not suitable for on-device learning scenarios. In practical on-device learning applications, data are collected in streaming. For example, the camera on the device keeps collecting video frames for hours or days, and the actions of nearly all classes are included in a single long video stream. Dividing such a long video stream into multiple video segments requires lots of human effort, which hinders the exploration of applying the TAL tasks to realistic on-device learning applications. To enable W-TAL models to learn from a long, untrimmed streaming video, we propose an efficient video learning approach that can directly adapt to new environments. We first propose a self-adaptive video dividing approach with a contrast score-based segment merging approach to convert the video stream into multiple segments. Then, we explore different sampling strategies on the TAL tasks to request as few labels as possible. To the best of our knowledge, we are the first attempt to directly learn from the on-device, long video stream. Experimental results on the THUMOS’14 dataset show that the performance of our approach is comparable to the current W-TAL state-of-the-art (SOTA) work without any laborious manual video splitting.