Multimodal Virtual Point 3D Detection

Multimodal Virtual Point 3D Detection
复制标题

DOI:
--
复制
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
通讯作者:
Tianwei Yin;Xingyi Zhou;Philipp Krähenbühl
Tianwei Yin;Xingyi Zhou;Philipp Krähenbühl
中科院分区:
其他
文献类型:
--
作者:
Tianwei Yin;Xingyi Zhou;Philipp Krähenbühl

文献摘要

相似文献

基于激光雷达的传感技术驱动着当前的自动驾驶汽车。尽管进展迅速,但目前的激光雷达传感器在分辨率和成本方面仍落后传统彩色相机20年。对于自动驾驶来说,这意味着靠近传感器的大物体很容易被看到,但远处或小物体只有一两个测量值。这是一个问题,尤其是当这些物体被证明是驾驶危险时。另一方面,这些相同的物体在机载RGB传感器上清晰可见。在这项工作中,我们提出了一种将RGB传感器无缝融合到基于激光雷达的3D识别中的方法。我们的方法采用一组2D检测来生成密集的3D虚拟点,以增强稀疏的3D点云。这些虚拟点自然集成到任何标准的基于激光雷达的3D探测器以及常规激光雷达测量。所得到的多模态检测器简单有效。在大规模nuScenes数据集上的实验结果表明,我们的框架提高了6.6 mAP的强中心点基线,并且优于竞争的融合方法。代码和更多的可视化可以在https://tianweiy.github.io/mvp/上获得
Lidar-based sensing drives current autonomous vehicles. Despite rapid progress, current Lidar sensors still lag two decades behind traditional color cameras in terms of resolution and cost. For autonomous driving, this means that large objects close to the sensors are easily visible, but far-away or small objects comprise only one measurement or two. This is an issue, especially when these objects turn out to be driving hazards. On the other hand, these same objects are clearly visible in onboard RGB sensors. In this work, we present an approach to seamlessly fuse RGB sensors into Lidar-based 3D recognition. Our approach takes a set of 2D detections to generate dense 3D virtual points to augment an otherwise sparse 3D point cloud. These virtual points naturally integrate into any standard Lidar-based 3D detectors along with regular Lidar measurements. The resulting multi-modal detector is simple and effective. Experimental results on the large-scale nuScenes dataset show that our framework improves a strong CenterPoint baseline by a significant 6.6 mAP, and outperforms competing fusion approaches. Code and more visualizations are available at https://tianweiy.github.io/mvp/