Compressed video ensemble based pseudo-labeling for semi-supervised action recognition

Compressed video ensemble based pseudo-labeling for semi-supervised action recognition
复制标题

DOI:
10.1016/j.mlwa.2022.100336
复制
发表时间:
2022-05
影响因子:
--
通讯作者:
Hayato Terao;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
Hayato Terao;Wataru Noguchi;H. Iizuka;Masahito Yamamoto
中科院分区:
--
文献类型:
--
作者:
Hayato Terao;Wataru Noguchi;H. Iizuka;Masahito Yamamoto

文献摘要

相似文献

最近的一些研究集中在基于深度学习的半监督学习的动作识别上。然而,很难扩大它们的训练,因为它们的输入是RGB帧,获得这些帧会产生计算和存储成本。在本文中,我们提出了一种半监督的动作识别方法,可以很容易地通过使用存储在压缩视频中的特征来扩大训练。我们的方法直接从压缩视频中提取多种类型的输入特征,而无需任何解码,并通过这些特征的预测组合来生成未标记视频的人工标签。除了对标记视频进行标准监督训练外,我们的模型还接受训练,以根据未标记压缩视频中的强增强特征预测人工标签。我们表明,我们的方法是更有效的,并实现了更好的分类性能在一些广泛使用的数据集比传统的半监督学习方法应用RGB帧。
Some recent studies have focused on deep learning based semi-supervised learning for action recognition. However, it is difficult to scale up their training because their input is RGB frames, the obtainment of which incurs computational and storage costs. In this paper, we propose a semi-supervised action recognition method that makes it easy to scale up the training by using features stored in compressed videos. Our method directly extracts multiple types of input features from compressed videos without any decoding and generates artificial labels of unlabeled videos through the ensembling of the predictions from these features. In addition to the standard supervised training on labeled videos, our models are trained to predict the artificial labels from strongly augmented features in unlabeled compressed videos. We show that our method is more efficient and achieves a better classification performance on some widely used datasets than conventional semi-supervised learning methods applying RGB frames.