Enabling spatio-temporal aggregation in Birds-Eye-View Vehicle Estimation

Enabling spatio-temporal aggregation in Birds-Eye-View Vehicle Estimation
复制标题

DOI:
10.1109/icra48506.2021.9561169
复制
发表时间:
2021-05
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Avishkar Saha;Oscar Alejandro Mendez Maldonado;Chris Russell;R. Bowden
Avishkar Saha;Oscar Alejandro Mendez Maldonado;Chris Russell;R. Bowden
中科院分区:
其他
文献类型:
--
作者:
Avishkar Saha;Oscar Alejandro Mendez Maldonado;Chris Russell;R. Bowden

文献摘要

相似文献

从单眼图像构建鸟瞰图通常是一个复杂的多阶段过程,涉及地平面估计、道路分割和三维目标检测等单独的视觉任务。然而,最近的方法采用了端到端解决方案,将基于图像的特征从图像平面扭曲到BEV,同时隐含地考虑了相机的几何形状。在这项工作中,我们展示了如何学习这种场景的瞬时BEV估计,并且可以通过结合时间信息来实现对世界的更好的状态估计。我们的模型通过分解3D卷积从单目视频中学习表征,并使用它来估计最终帧的BEV占用网格。我们获得了最先进的单眼图像BEV估计结果,并建立了单眼视频单场景BEV估计的新基准。
Constructing Birds-Eye-View (BEV) maps from monocular images is typically a complex multi-stage process involving the separate vision tasks of ground plane estimation, road segmentation and 3D object detection. However, recent approaches have adopted end-to-end solutions which warp image-based features from the image-plane to BEV while implicitly taking account of camera geometry. In this work, we show how such instantaneous BEV estimation of a scene can be learnt, and a better state estimation of the world can be achieved by incorporating temporal information. Our model learns a representation from monocular video through factorised 3D convolutions and uses this to estimate a BEV occupancy grid of the final frame. We achieve state-of-the-art results for BEV estimation from monocular images, and establish a new benchmark for single-scene BEV estimation from monocular video.