Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes

Augmented Reality Meets Computer Vision: Efficient Data Generation for Urban Driving Scenes
复制标题

DOI:
10.1007/s11263-018-1070-x
复制
发表时间:
2018-09-01
影响因子:
19.5
通讯作者:
Rother, Carsten
Rother, Carsten
中科院分区:
计算机科学2区
文献类型:
--
作者:
Abu Alhaija, Hassan;Mustikovela, Siva Karthik;Rother, Carsten

文献摘要

被引文献

相似文献

深度学习在计算机视觉中的成功是基于大型注释数据集的可用性。为了降低对手工标记图像的需求,虚拟渲染的3D世界最近越来越受欢迎。不幸的是,创建逼真的3D内容本身就具有挑战性,需要大量的人力。在这项工作中,我们提出了一种替代的范式,结合了真实的和合成数据学习语义实例分割和对象检测模型。利用并非场景的所有方面对于此任务都同样重要的事实,我们建议使用目标类别的虚拟对象来增强真实世界图像。捕获大规模的真实世界图像容易且便宜,并且直接提供真实的背景外观,而不需要创建环境的复杂3D模型。我们提出了一个有效的程序来增强这些图像与虚拟对象。与建模完整的3D环境相比,我们的数据增强方法只需要一些用户交互与目标对象类别的3D模型相结合。利用我们的方法,我们引入了一个新的增强城市驾驶场景数据集,其中包含360度图像,这些图像用作环境地图,以在渲染对象上创建逼真的照明和反射。我们分析了现实的对象放置的意义,通过比较人工放置人类的语义场景分析的基础上的自动方法。这使我们能够创建既表现出现实的背景外观,以及大量的复杂的对象安排的合成图像。通过大量的实验,我们得出了正确的参数集,以产生增强的数据,可以最大限度地提高实例分割模型的性能。此外,我们展示了所提出的方法在训练标准深度模型以用于室外驾驶场景中的汽车的语义实例分割和对象检测方面的实用性。我们在KITTI 2015数据集和Cityscapes数据集上测试了在我们的增强数据上训练的模型,我们已经用像素精确的地面实况进行了注释。我们的实验表明,在增强图像上训练的模型比在完全合成的数据上训练的模型或在有限数量的带注释的真实的数据上训练的模型更好地泛化。
The success of deep learning in computer vision is based on the availability of large annotated datasets. To lower the need for hand labeled images, virtually rendered 3D worlds have recently gained popularity. Unfortunately, creating realistic 3D content is challenging on its own and requires significant human effort. In this work, we propose an alternative paradigm which combines real and synthetic data for learning semantic instance segmentation and object detection models. Exploiting the fact that not all aspects of the scene are equally important for this task, we propose to augment real-world imagery with virtual objects of the target category. Capturing real-world images at large scale is easy and cheap, and directly provides real background appearances without the need for creating complex 3D models of the environment. We present an efficient procedure to augment these images with virtual objects. In contrast to modeling complete 3D environments, our data augmentation approach requires only a few user interactions in combination with 3D models of the target object category. Leveraging our approach, we introduce a novel dataset of augmented urban driving scenes with 360 degree images that are used as environment maps to create realistic lighting and reflections on rendered objects. We analyze the significance of realistic object placement by comparing manual placement by humans to automatic methods based on semantic scene analysis. This allows us to create composite images which exhibit both realistic background appearance as well as a large number of complex object arrangements. Through an extensive set of experiments, we conclude the right set of parameters to produce augmented data which can maximally enhance the performance of instance segmentation models. Further, we demonstrate the utility of the proposed approach on training standard deep models for semantic instance segmentation and object detection of cars in outdoor driving scenarios. We test the models trained on our augmented data on the KITTI 2015 dataset, which we have annotated with pixel-accurate ground truth, and on the Cityscapes dataset. Our experiments demonstrate that the models trained on augmented imagery generalize better than those trained on fully synthetic data or models trained on limited amounts of annotated real data.