Bidirectional LSTM with saliency-aware 3D-CNN features for human action recognition

Bidirectional LSTM with saliency-aware 3D-CNN features for human action recognition
复制标题

DOI:
10.36909/jer.v9i3a.8383
复制
发表时间:
2021-09-01
影响因子:
1
通讯作者:
Hussain, Fida
Hussain, Fida
中科院分区:
工程技术4区
文献类型:
--
作者:
Arif, Sheeraz;Wang, Jing;Hussain, Fida

文献摘要

被引文献

相似文献

深度卷积神经网络(DCNN)和递归神经网络(RNN)已被证明是多媒体理解中的一个迫切的研究领域,并取得了显着的动作识别性能。然而,视频包含具有不同维度的丰富运动信息。现有的基于循环的流水线无法捕获具有各种运动尺度和由多个演员执行的复杂动作的视频中的长期运动动态。考虑上下文和显著特征比将视频帧映射到静态视频表示更重要。这项研究工作通过使用3D卷积(C3D)网络和新引入的深度双向LSTM分析和处理视频信息提供了一种新的管道。像流行的双流修道院,我们也引入了一个双流框架的修改,即我们取代的光流流显着性感知流,以避免计算复杂度。首先,我们通过应用显著性感知方法来生成显著性感知视频流。其次,双流3D卷积网络(C3D)被用于两种不同类型的流,即,RGB流和显著性感知视频流,以收集空间和语义时间特征。接下来,使用深度双向LSTM网络来学习序列深度时间动态。最后,时间序列池层和softmax层对人类活动和行为进行分类。引入的系统可以学习长期的时间依赖性,并可以预测复杂的人类行为。实验结果表明,在不同的基准数据集上的动作识别精度显着提高。
Deep convolutional neural network (DCNN) and recurrent neural network (RNN) have been proved as an imperious research area in multimedia understanding and obtained remarkable action recognition performance. However, videos contain rich motion information with varying dimensions. Existing recurrent based pipelines fail to capture long-term motion dynamics in videos with various motion scales and complex actions performed by multiple actors. Consideration of contextual and salient features is more important than mapping a video frame into a static video representation. This research work provides a novel pipeline by analyzing and processing the video information using a 3D convolution (C3D) network and newly introduced deep bidirectional LSTM. Like popular two-stream convent, we also introduce a two-stream framework with one modification; that is, we replace the optical flow stream by saliency-aware stream to avoid the computational complexity. First, we generate a saliency-aware video stream by applying the saliency-aware method. Secondly, a two-stream 3D-convolutional network (C3D) is utilized with two different types of streams, i.e., RGB stream and saliency-aware video stream, to collect both spatial and semantic temporal features. Next, a deep bidirectional LSTM network is used to learn sequential deep temporal dynamics Finally, time-series-pooling-layer and softmax-layers classify human activity and behavior. The introduced system can learn long-term temporal dependencies and can predict complex human actions. Experimental results demonstrate the significant improvement in action recognition accuracy on different benchmark datasets.