Evaluating fusion of RGB-D and inertial sensors for multimodal human action recognition

Evaluating fusion of RGB-D and inertial sensors for multimodal human action recognition
复制标题

DOI:
10.1007/s12652-019-01239-9
复制
发表时间:
2020-01-01
影响因子:
--
通讯作者:
Raman, Balasubramanian
Raman, Balasubramanian
中科院分区:
计算机科学3区
文献类型:
--
作者:
Imran, Javed;Raman, Balasubramanian

文献摘要

被引文献

相似文献

融合来自不同传感器的多模态是多模态人体动作识别的一个重要研究领域。在本文中,我们进行了深入的研究,以调查不同的参数,如输入预处理,数据增强,网络架构和模型融合的影响,以便提出一个实用的指导方针,多模态动作识别使用深度学习范式。首先,对于RGB视频,我们提出了一种新的基于图像的描述符称为堆叠的稠密流差分图像(SDFDI),能够捕获视频序列中存在的时空信息。然后训练各种深度2D卷积神经网络(CNN),将我们的SDFDI与最先进的基于图像的表示进行比较。其次,对于骨架流,我们提出了基于3D变换的数据增强技术,以便于在小数据集上训练深度神经网络。我们还提出了一个基于双向门控递归单元(BiGRU)的递归神经网络(RNN)来建模骨架数据。第三,对于惯性传感器数据,我们提出了基于白色高斯噪声抖动的数据增强,沿着深度1D-CNN网络进行动作分类。所有这三种异构网络(1D-CNN,2D-CNN和BiGRU)的输出通过基于得分和特征融合的各种模型融合方法进行组合。最后,为了说明所提出的框架的有效性,我们在公开可用的UTD-MHAD数据集上测试了我们的模型,并实现了97.91%的总体准确率,比单独使用每种模态高出约4%。希望本文的讨论和结论能为相关领域的研究人员提供更深入的见解,并为不同的多传感器融合体系结构的进一步研究提供思路。
Fusion of multiple modalities from different sensors is an important area of research for multimodal human action recognition. In this paper, we conduct an in-depth study to investigate the effect of different parameters like input preprocessing, data augmentation, network architectures and model fusion so as to come up with a practical guideline for multimodal action recognition using deep learning paradigm. First, for RGB videos, we propose a novel image-based descriptor called stacked dense flow difference image (SDFDI), capable of capturing the spatio-temporal information present in a video sequence. A variety of deep 2D convolutional neural networks (CNN) are then trained to compare our SDFDI against state-of-the-art image-based representations. Second, for skeleton stream, we propose data augmentation technique based on 3D transformations so as to facilitate training a deep neural network on small datasets. We also propose a bidirectional gated recurrent unit (BiGRU) based recurrent neural network (RNN) to model skeleton data. Third, for inertial sensor data, we propose data augmentation based on jittering with white Gaussian noise along with deep a 1D-CNN network for action classification. The outputs of all these three heterogeneous networks (1D-CNN, 2D-CNN and BiGRU) are combined by a variety of model fusion approach based on score and feature fusion. Finally, in order to illustrate the efficacy of the proposed framework, we test our model on a publicly available UTD-MHAD dataset, and achieved an overall accuracy of 97.91%, which is about 4% higher than using each modality individually. We hope that the discussions and conclusions from this work will provide a deeper insight to the researchers in the related fields, and provide avenues for further studies for different multi-sensor based fusion architectures.