RGB-D-Based Object Recognition Using Multimodal Convolutional Neural Networks: A Survey

RGB-D-Based Object Recognition Using Multimodal Convolutional Neural Networks: A Survey
复制标题

使用多模态卷积神经网络进行基于 RGB-D 的物体识别:一项调查

DOI:
10.1109/access.2019.2907071
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Liu, Zheng
Liu, Zheng
中科院分区:
计算机科学3区
文献类型:
--
作者:
Gao, Mingliang;Jiang, Jun;Liu, Zheng

文献摘要

被引文献

相似文献

现实环境中的物体识别是计算机视觉和机器人领域的基本和关键任务之一。借助先进的传感技术和低成本的深度传感器,可以同步记录高质量的RGB和深度图像,并通过联合利用它们来提高对象识别性能。基于RGB-D的对象识别已经从早期使用手工制作的表示方法发展到目前最先进的基于深度学习的方法。随着深度学习,特别是卷积神经网络(CNN)在视觉领域取得的不可否认的成功,深度学习研究的自然进展指向涉及更大,更复杂的多模态数据的问题。在本文中,我们提供了一个全面的调查,最近的多模态CNN(MMCNN)为基础的方法,已经证明了显着的改进,比以前的方法。我们强调两个关键问题,即训练数据不足和多模态融合。此外,我们总结并讨论了公开可用的RGB-D对象识别数据集,并在这些基准数据集上对所提出的方法进行了比较性能评估。最后,我们确定在这个迅速发展的领域有前途的研究途径。这项调查不仅使研究人员能够很好地概述基于RGB-D的对象识别的最新方法,而且还为其他多模态机器学习应用提供参考,例如,多模态医学图像融合、视听语音识别以及多媒体检索和生成。
Object recognition in real-world environments is one of the fundamental and key tasks in computer vision and robotics communities. With the advanced sensing technologies and low-cost depth sensors, the high-quality RGB and depth images can be recorded synchronously, and the object recognition performance can be improved by jointly exploiting them. RGB-D-based object recognition has evolved from early methods that using hand-crafted representations to the current state-of-the-art deep learning-based methods. With the undeniable success of deep learning, especially convolutional neural networks (CNNs) in the visual domain, the natural progression of deep learning research points to problems involving larger and more complex multimodal data. In this paper, we provide a comprehensive survey of recent multimodal CNNs (MMCNNs)-based approaches that have demonstrated significant improvements over previous methods. We highlight two key issues, namely, training data deficiency and multimodal fusion. In addition, we summarize and discuss the publicly available RGB-D object recognition datasets and present a comparative performance evaluation of the proposed methods on these benchmark datasets. Finally, we identify promising avenues of research in this rapidly evolving field. This survey will not only enable researchers to get a good overview of the state-of-the-art methods for RGB-D-based object recognition but also provide a reference for other multimodal machine learning applications, e.g., multimodal medical image fusion, audio-visual speech recognition, and multimedia retrieval and generation.