Weakly-Supervised Action Localization, and Action Recognition Using Global-Local Attention of 3D CNN

Weakly-Supervised Action Localization, and Action Recognition Using Global-Local Attention of 3D CNN
复制标题

弱监督动作定位和使用 3D CNN 全局局部注意力的动作识别

DOI:
10.1007/s11263-022-01649-x
复制
发表时间:
2022
影响因子:
19.5
通讯作者:
Kurita Takio
Kurita Takio
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yudistira Novanto;Kavitha Muthu Subash;Kurita Takio

文献摘要

相似文献

3D卷积神经网络(3D CNN)捕获3D数据(如视频序列)的空间和时间信息。然而,由于卷积和池化机制,发生的信息丢失似乎是不可避免的。为了改善3D CNN中的视觉解释和分类,我们提出了两种方法;(i)使用训练好的3DResNext网络聚合逐层全局到局部(全局-局部)离散梯度,以及(ii)实现注意力门控网络以提高动作识别的准确性。所提出的方法旨在通过视觉归因,弱监督动作定位和动作识别来显示3D CNN中称为全局-局部注意力的每一层的有用性。首先,3DResNext被训练并应用于使用关于最大预测类的反向传播的动作分类。然后对每个层的梯度和激活进行上采样。之后,聚合被用来产生更细微的注意力,指出预测类的输入视频中最关键的部分。我们使用轮廓阈值的最终注意力的最终定位。我们通过3DCAM使用细粒度的视觉解释来评估修剪视频中的空间和时间动作定位。实验结果表明,该方法产生信息丰富的视觉解释和歧视性的注意。此外,通过每层的注意门控动作识别产生更好的分类结果比基线模型。
3D convolutional neural network (3D CNN) captures spatial and temporal information on 3D data such as video sequences. However, due to the convolution and pooling mechanism, the information loss that occurs seems unavoidable. To improve the visual explanations and classification in 3D CNN, we propose two approaches; (i) aggregate layer-wise global to local (global–local) discrete gradient using trained 3DResNext network, and (ii) implement attention gating network to improve the accuracy of the action recognition. The proposed approach intends to show the usefulness of every layer termed as global–local attention in 3D CNN via visual attribution, weakly-supervised action localization, and action recognition. Firstly, the 3DResNext is trained and applied for action classification using backpropagation concerning the maximum predicted class. The gradient and activation of every layer are then up-sampled. Later, aggregation is used to produce more nuanced attention, which points out the most critical part of the predicted class’s input videos. We use contour thresholding of final attention for final localization. We evaluate spatial and temporal action localization in trimmed videos using fine-grained visual explanation via 3DCAM. Experimental results show that the proposed approach produces informative visual explanations and discriminative attention. Furthermore, the action recognition via attention gating of each layer produces better classification results than the baseline model.