MonoFENet: Monocular 3D Object Detection With Feature Enhancement Networks

MonoFENet: Monocular 3D Object Detection With Feature Enhancement Networks
复制标题

MonoFENet:具有特征增强网络的单目 3D 物体检测

DOI:
10.1109/tip.2019.2952201
复制
发表时间:
2020-01-01
影响因子:
10.6
通讯作者:
Chen, Zhenzhong
Chen, Zhenzhong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bao, Wentao;Xu, Bin;Chen, Zhenzhong

文献摘要

被引文献

相似文献

单目三维目标检测具有成本低、可作为自动驾驶系统辅助模块的优点,近年来受到越来越多的关注。本文提出了一种基于特征增强网络的单目三维目标检测方法,我们称之为MonoFENet。具体来说,通过输入单眼图像估计的视差,可以增强二维和三维流的特征,并利用其进行精确的三维定位。对于二维流,输入图像用于生成二维区域建议以及提取外观特征。对于三维流,估计的视差被转换成三维密集点云,然后通过相关的前视图地图增强。在RoI Mean Pooling层的基础上,提出的点特征增强(PointFE)网络进一步增强了RoI点云的三维几何特征。融合图像和点云的区域特征进行最终的二维和三维边界盒回归。在KITTI基准上的实验结果表明,我们的方法可以达到最先进的单眼三维目标检测性能。
Monocular 3D object detection has the merit of low cost and can be served as an auxiliary module for autonomous driving system, becoming a growing concern in recent years. In this paper, we present a monocular 3D object detection method with feature enhancement networks, which we call MonoFENet. Specifically, with the estimated disparity from the input monocular image, the features of both the 2D and 3D streams can be enhanced and utilized for accurate 3D localization. For the 2D stream, the input image is used to generate 2D region proposals as well as to extract appearance features. For the 3D stream, the estimated disparity is transformed into 3D dense point cloud, which is then enhanced by the associated front view maps. With the RoI Mean Pooling layer, 3D geometric features of RoI point clouds are further enhanced by the proposed point feature enhancement (PointFE) network. The region-wise features of image and point cloud are fused for the final 2D and 3D bounding boxes regression. The experimental results on the KITTI benchmark reveal that our method can achieve state-of-the-art performance for monocular 3D object detection.