Recognizing Actions in Videos from Unseen Viewpoints

Recognizing Actions in Videos from Unseen Viewpoints
复制标题

DOI:
10.1109/cvpr46437.2021.00411
复制
发表时间:
2021-03
期刊:
2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
A. Piergiovanni;M. Ryoo
A. Piergiovanni;M. Ryoo
中科院分区:
其他
文献类型:
--
作者:
A. Piergiovanni;M. Ryoo

文献摘要

相似文献

视频识别的标准方法使用旨在捕获时空数据的大型CNN。然而,训练这些模型需要大量的标签训练数据,包含各种各样的动作、场景、设置和摄像机视点。在本文中,我们证明了现有的卷积神经网络模型不能识别其训练数据中不存在的摄像机视点的动作(即看不见的视点动作识别)。为了解决这一问题,我们发展了基于3D表示的方法,并引入了一个新的几何卷积层来学习视点不变表示。此外,我们介绍了一个新的,具有挑战性的数据集,用于不可见的视图识别,并展示了该方法学习视点不变表示的能力。
Standard methods for video recognition use large CNNs designed to capture spatio-temporal data. However, training these models requires a large amount of labeled training data, containing a wide variety of actions, scenes, settings and camera viewpoints. In this paper, we show that current convolutional neural network models are unable to recognize actions from camera viewpoints not present in their training data (i.e., unseen view action recognition). To address this, we develop approaches based on 3D representations and introduce a new geometric convolutional layer that can learn viewpoint invariant representations. Further, we introduce a new, challenging dataset for unseen view recognition and show the approaches ability to learn viewpoint invariant representations.