Image-to-Image Translation for Autonomous Driving from Coarsely-Aligned Image Pairs

Image-to-Image Translation for Autonomous Driving from Coarsely-Aligned Image Pairs
复制标题

DOI:
10.1109/icra48891.2023.10160815
复制
发表时间:
2022-09
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Youya Xia;Josephine Monica;Wei-Lun Chao;Bharath Hariharan;Kilian Q. Weinberger;Mark E. Campbell
Youya Xia;Josephine Monica;Wei-Lun Chao;Bharath Hariharan;Kilian Q. Weinberger;Mark E. Campbell
中科院分区:
其他
文献类型:
--
作者:
Youya Xia;Josephine Monica;Wei-Lun Chao;Bharath Hariharan;Kilian Q. Weinberger;Mark E. Campbell

文献摘要

被引文献

相似文献

自动驾驶汽车必须能够可靠地应对恶劣的天气条件(例如,(雪)安全运行。在本文中,我们研究了转动传感器输入的想法(即,图像)转换为良性(即,晴天),在其上下游任务(例如,语义分割)可以达到高精度。先前的工作主要将其表述为未配对的图像到图像的翻译问题,这是由于缺乏在完全相同的相机姿势和语义布局下捕获的配对图像。虽然完美对齐的图像不可用,但可以容易地获得粗略配对的图像。例如,许多人每天在好天气和坏天气下驾驶相同的路线;因此,在附近的GPS位置捕获的图像可以形成一对。虽然来自重复遍历的数据不太可能捕获相同的前景对象,但我们认为它们提供了丰富的上下文信息来监督图像翻译模型。为此,我们提出了一种新的训练目标,利用粗对齐的图像对。我们表明,我们的粗对齐训练方案导致更好的图像翻译质量和改进的下游任务,如语义分割,单目深度估计和视觉定位。
A self-driving car must be able to reliably handle adverse weather conditions (e.g., snowy) to operate safely. In this paper, we investigate the idea of turning sensor inputs (i.e., images) captured in an adverse condition into a benign one (i.e., sunny), upon which the downstream tasks (e.g., semantic segmentation) can attain high accuracy. Prior work primarily formulates this as an unpaired image-to-image translation problem due to the lack of paired images captured under the exact same camera poses and semantic layouts. While perfectly-aligned images are not available, one can easily obtain coarsely-paired images. For instance, many people drive the same routes daily in both good and adverse weather; thus, images captured at close-by GPS locations can form a pair. Though data from repeated traversals are unlikely to capture the same foreground objects, we posit that they provide rich contextual information to supervise the image translation model. To this end, we propose a novel training objective leveraging coarsely-aligned image pairs. We show that our coarsely-aligned training scheme leads to a better image translation quality and improved downstream tasks, such as semantic segmentation, monocular depth estimation, and visual localization.