Free Supervision from Video Games

Free Supervision from Video Games
复制标题

DOI:
10.1109/cvpr.2018.00312
复制
发表时间:
2018-06
期刊:
2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition
影响因子:
--
通讯作者:
Philipp Krähenbühl
Philipp Krähenbühl
中科院分区:
其他
文献类型:
--
作者:
Philipp Krähenbühl

文献摘要

被引文献

相似文献

深度网络非常需要数据。他们吞噬了成千上万的标记图像来学习鲁棒和语义上有意义的特征表示。当前的网络非常需要数据,因此收集标签数据变得和设计网络本身一样重要。不幸的是,手动数据收集既昂贵又耗时。我们提出了另一种选择,并展示了如何在我们玩视频游戏时实时从视频游戏中轻松提取许多视觉任务的真实标签。我们为流行的Microsoft®DirectX®渲染API提供接口,并在游戏运行时注入专门的渲染代码。此代码为实例分割、语义标记、深度估计、光流、固有图像分解和实例跟踪生成地面真值标签。现在,研究人员不再给图像贴上标签,而是整天玩电子游戏。我们的方法是通用的,适用于各种电子游戏。我们收集了一个包含220k训练图像和60k测试图像的数据集,并评估了最先进的光流、深度估计和内在图像分解算法。我们的电子游戏数据在视觉上比其他合成数据集更接近真实世界的图像。
Deep networks are extremely hungry for data. They devour hundreds of thousands of labeled images to learn robust and semantically meaningful feature representations. Current networks are so data hungry that collecting labeled data has become as important as designing the networks themselves. Unfortunately, manual data collection is both expensive and time consuming. We present an alternative, and show how ground truth labels for many vision tasks are easily extracted from video games in real time as we play them. We interface the popular Microsoft® DirectX® rendering API, and inject specialized rendering code into the game as it is running. This code produces ground truth labels for instance segmentation, semantic labeling, depth estimation, optical flow, intrinsic image decomposition, and instance tracking. Instead of labeling images, a researcher now simply plays video games all day long. Our method is general and works on a wide range of video games. We collected a dataset of 220k training images, and 60k test images across 3 video games, and evaluate state of the art optical flow, depth estimation and intrinsic image decomposition algorithms. Our video game data is visually closer to real world images, than other synthetic dataset.