Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised Approach

Extrinsic Self-Calibration of the Surround-View System: A Weakly Supervised Approach
复制标题

DOI:
10.1109/tmm.2022.3144889
复制
发表时间:
2023
影响因子:
7.3
通讯作者:
Yang Chen;Lin Zhang;Ying Shen;Brian Nlong Zhao;Yicong Zhou
Yang Chen;Lin Zhang;Ying Shen;Brian Nlong Zhao;Yicong Zhou
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yang Chen;Lin Zhang;Ying Shen;Brian Nlong Zhao;Yicong Zhou

文献摘要

相似文献

SVS通常由四个安装在车辆周围的广角鱼眼摄像头组成,以感知周围环境。在标定摄像机内外部参数的前提下,利用摄像机同步采集的图像可以合成自顶向下的全景图。目前,内部校准方法已经比较完善,可以流水线化,而外部校准方法还不成熟。为了填补这一研究空白,我们提出了一种新的外部自校准方案,它遵循一个弱监督的框架,即WESNet(弱监督外部自校准网络)。WESNet的培训包括两个阶段。首先,我们利用几个校准站点图像中的角点作为弱监督,通过最小化几何损失来粗略优化网络。然后,在第一阶段的收敛之后,我们另外引入了一个自监督的光度损失项,该损失项可以通过来自自然图像的光度信息来构造,以进行进一步的微调。此外,为了支持训练,我们共收集了19,078组在各种环境条件下同步捕获的鱼眼图像。据我们所知,这是迄今为止包含原始鱼眼图像的最大环绕视图数据集。通过从训练数据中学习先验知识,WESNet将同步采集的原始鱼眼图像作为输入,直接产生端到端的外部特征,只需很少的人工成本。它的效率和功效已经通过对我们收集的数据集进行的大量实验得到了证实。为了使我们的结果具有可重复性,源代码和收集的数据集已经发布。
An SVS usually consists of four wide-angle fisheye cameras mounted around the vehicle to sense the surrounding environment. From the images synchronously captured by cameras, a top-down surround-view can be synthesized, on the premise that both intrinsics and extrinsics of the cameras have been calibrated. At present, the intrinsic calibration approach is relatively complete and can be pipelined, while the extrinsic calibration is still immature. To fill such a research gap, we propose a novel extrinsic self-calibration scheme which follows a weakly supervised framework, namely WESNet (Weakly-supervised Extrinsic Self-calibration Network). The training of WESNet consists of two stages. First, we utilize the corners in a few calibration site images as the weak supervision to roughly optimize the network by minimizing the geometric loss. Then, after the convergence in the first stage, we additionally introduce a self-supervised photometric loss term that can be constructed by the photometric information from natural images for further fine-tuning. Besides, to support training, we totally collected 19,078 groups of synchronously captured fisheye images under various environmental conditions. To our knowledge, thus far this is the largest surround-view dataset containing original fisheye images. By means of learning prior knowledge from the training data, WESNet takes the original fisheye images synchronously collected as the input, and directly yields extrinsics end-to-end with little labor cost. Its efficiency and efficacy have been corroborated by extensive experiments conducted on our collected dataset. To make our results reproducible, source code and the collected dataset have been released.1