Spatiotemporal Initialization for 3D CNNs with Generated Motion Patterns

Spatiotemporal Initialization for 3D CNNs with Generated Motion Patterns
复制标题

DOI:
10.1109/wacv51458.2022.00081
复制
发表时间:
2022-01
期刊:
2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Hirokatsu Kataoka;Kensho Hara;Ryusuke Hayashi;Eisuke Yamagata;Nakamasa Inoue
Hirokatsu Kataoka;Kensho Hara;Ryusuke Hayashi;Eisuke Yamagata;Nakamasa Inoue
中科院分区:
其他
文献类型:
--
作者:
Hirokatsu Kataoka;Kensho Hara;Ryusuke Hayashi;Eisuke Yamagata;Nakamasa Inoue

文献摘要

相似文献

该论文提出了一种用于时空初始化的公式驱动监督学习(FDSL)框架。我们的 FDSL 方法能够使用基于 Perlin 噪声的简单公式自动同时生成运动模式及其视频标签。我们设计了一个生成的运动模式数据集,足以让 3D CNN 学习更好的自然视频基础集。构建的视频柏林噪声(VPN)数据集可用于在使用 Kinetics-400/700 等大规模视频数据集进行预训练之前初始化模型,以提高目标任务性能。我们使用 VPN 数据集的时空初始化(VPN 初始化)优于之前使用 2D ImageNet 数据集的膨胀 3D ConvNet (I3D) 的初始化方法。我们提出的方法提高了 Kinetics-400 预训练模型在 {Kinetics-400、UCF-101、HMDB-51、ActivityNet} 数据集上的 top-1 视频级精度。特别是,所提出的方法将 Kinetics-400 预训练模型在 ActivityNet 上的性能提高了 10.3 pt。我们还报告说,3D CNN 相对于基线的相对性能改进比其他模型更大。我们的 VPN 初始化主要有助于增强时空 3D 内核的性能。本研究中使用的数据集、代码和预训练模型将公开可用1。
The paper proposes a framework of Formula-Driven Supervised Learning (FDSL) for spatiotemporal initialization. Our FDSL approach enables to automatically and simultaneously generate motion patterns and their video labels with a simple formula which is based on Perlin noise. We designed a dataset of generated motion patterns adequate for the 3D CNNs to learn a better basis set of natural videos. The constructed Video Perlin Noise (VPN) dataset can be applied to initialize a model before pre-training with large-scale video datasets such as Kinetics-400/700, to enhance target task performance. Our spatiotemporal initialization with VPN dataset (VPN initialization) outperforms the previous initialization method with the inflated 3D ConvNet (I3D) using 2D ImageNet dataset. Our proposed method increased the top-1 video-level accuracy of Kinetics-400 pre-trained model on {Kinetics-400, UCF-101, HMDB-51, ActivityNet} datasets. Especially, the proposed method increased the performance rate of Kinetics-400 pre-trained model by 10.3 pt on ActivityNet. We also report that the relative performance improvements from the baseline are greater in 3D CNNs rather than other models. Our VPN initialization mainly helps to enhance the performance in spatiotemporal 3D kernels. The datasets, codes and pre-trained models used in this study will be publicly available1.