3D-SSD: Learning hierarchical features from RGB-D images for amodal 3D object detection

3D-SSD: Learning hierarchical features from RGB-D images for amodal 3D object detection
复制标题

DOI:
10.1016/j.neucom.2019.10.025
复制
发表时间:
2017-11
期刊:
影响因子:
6
通讯作者:
Qianhui Luo;Huifang Ma;Li Tang;Yue Wang;R. Xiong
Qianhui Luo;Huifang Ma;Li Tang;Yue Wang;R. Xiong
中科院分区:
计算机科学2区
文献类型:
--
作者:
Qianhui Luo;Huifang Ma;Li Tang;Yue Wang;R. Xiong

文献摘要

被引文献

相似文献

本文旨在开发一种更快,更准确的解决方案,非模态的室内场景的三维物体检测问题。该解决方案是通过一种新的神经网络结构来实现的,该结构将一对RGB-D图像作为输入,并将定向的3D边界框作为输出。这种网络称为3D-SSD,有两个组成部分:分层特征融合和多层预测。分层特征融合结合了从RGB-D图像中学习到的多尺度外观和几何特征,这些特征随后用于目标检测的多层预测。通过以协同的方式利用2.5D表示,可以提高精度和效率。为了专门解决不同对象的形状差异,将一组具有不同物理尺寸的3D锚框附加到预测层上的每个位置。在测试时,3D锚框的类别分数是通过调整位置、大小和方向生成的,从而导致使用非最大抑制的最终检测。在SUN RGB-D和NYUV 2的公开数据集上进行了全面的实验。结果表明,该算法是第一个在具有挑战性的数据集上近实时运行的3D检测器,与最先进的方法相比具有竞争力的性能。3D-SSD在SUN RGB-D数据集上以约5.6 fps的速度获得37.1% mAP,比最先进的Deep Sliding Shape快10.2% mAP,快约109倍。对于速率为9.3 fps的高效模型设置,3D-SSD在mAP上的准确率仍为37%。此外,实验还表明,所提出的方法实现了相当的准确性,即使在输入图像尺寸较小的情况下,也比NYUv 2数据集上的最新方法快约477倍。
This paper aims at developing a faster and more accurate solution to the amodal 3D object detection problem for indoor scenarios. The solution is achieved through a novel neural network structure which takes a pair of RGB-D images as input and delivers oriented 3D bounding boxes as the output. Such network, named 3D-SSD, has two components: hierarchical feature fusion and multi-layer prediction. The hierarchical feature fusion combines multi-scale appearance and geometric features learned from RGB-D images, which is later utilized in the multi-layer prediction for object detection. Both the accuracy and the efficiency can be improved by exploiting 2.5D representations in a synergistic way. To specifically address the shape variance of different objects, a set of 3D anchor boxes with varying physical sizes are attached to every location on the prediction layers. While testing, the category scores for 3D anchor boxes are generated with adjusted positions, sizes and orientations, leading to the final detections using non-maximum suppression. Comprehensive experiments have been performed on publicly accessible dataset of SUN RGB-D and NYUV2. The results show the proposed algorithm is the first 3D detector that runs in near real-time on the challenging datasets with competitive performance to the state-of-the-art methods. The 3D-SSD gets 37.1% mAP on the SUN RGB-D dataset at around 5.6 fps, which outperforms the state-of-the-art Deep Sliding Shape by 10.2% mAP and around 109 ×  faster. For an efficient model setting with a rate of 9.3 fps, 3D-SSD still gets an accuracy of 37% on mAP. Further, experiments also suggest the proposed approach achieves comparable accuracy and is about 477 ×  faster than the state-of-art method on the NYUv2 dataset even with a smaller input image size.