Early vs Late Fusion in Multimodal Convolutional Neural Networks

Early vs Late Fusion in Multimodal Convolutional Neural Networks
复制标题

DOI:
10.23919/fusion45008.2020.9190246
复制
发表时间:
2020-07
期刊:
2020 IEEE 23rd International Conference on Information Fusion (FUSION)
影响因子:
--
通讯作者:
K. Gadzicki;Razieh Khamsehashari;C. Zetzsche
K. Gadzicki;Razieh Khamsehashari;C. Zetzsche
中科院分区:
其他
文献类型:
--
作者:
K. Gadzicki;Razieh Khamsehashari;C. Zetzsche

文献摘要

被引文献

相似文献

将神经网络中的机器学习与多模态融合策略相结合,为分类任务提供了一个有趣的潜力,但许多应用的最佳融合策略尚未确定。在这里,我们在人类活动识别的背景下解决这个问题,利用最先进的卷积网络架构(Inception I3D)和巨大的数据集(NTU RGB+D)。作为模态,我们考虑RGB视频,光流和骨架数据。我们确定不同模态的融合是否可以提供一个优势相比,单模态的方法,以及是否一个更复杂的早期融合策略可以通过利用不同模态之间的统计相关性优于简单的后期融合策略。我们的研究结果表明,通过多模态融合和早期融合策略的实质性优势,
Combining machine learning in neural networks with multimodal fusion strategies offers an interesting potential for classification tasks but the optimum fusion strategies for many applications have yet to be determined. Here we address this issue in the context of human activity recognition, making use of a state-of-the-art convolutional network architecture (Inception I3D) and a huge dataset (NTU RGB+D). As modalities we consider RGB video, optical flow, and skeleton data. We determine whether the fusion of different modalities can provide an advantage as compared to uni-modal approaches, and whether a more complex early fusion strategy can outperform the simpler late-fusion strategy by making use of statistical correlations between the different modalities. Our results show a clear performance improvement by multi-modal fusion and a substantial advantage of an early fusion strategy,