EventGAN: Leveraging Large Scale Image Datasets for Event Cameras

EventGAN: Leveraging Large Scale Image Datasets for Event Cameras
复制标题

DOI:
10.1109/iccp51581.2021.9466265
复制
发表时间:
2019-12
期刊:
2021 IEEE International Conference on Computational Photography (ICCP)
影响因子:
--
通讯作者:
A. Z. Zhu;ZiYun Wang;Kaung Khant;Kostas Daniilidis
A. Z. Zhu;ZiYun Wang;Kaung Khant;Kostas Daniilidis
中科院分区:
其他
文献类型:
--
作者:
A. Z. Zhu;ZiYun Wang;Kaung Khant;Kostas Daniilidis

文献摘要

被引文献

相似文献

与传统相机相比,活动相机具有许多优点,例如能够跟踪令人难以置信的快速运动,高动态范围和低功耗。然而,它们在计算机视觉问题中的应用,其中许多主要由深度学习解决方案主导,由于缺乏事件的标记训练数据而受到限制。在这项工作中,我们提出了一种方法,该方法通过使用卷积神经网络模拟来自一对时间图像帧的事件来利用图像的现有标记数据。我们在成对的图像和事件上训练这个网络,使用一个对抗性的一致性损失和一对周期一致性损失。循环一致性损失利用一对预训练的自监督网络,该网络从事件执行光流估计和图像重建,并约束我们的网络生成事件,从而从这两个网络中获得准确的输出。经过完全端到端的训练,我们的网络从图像中学习事件的生成模型,而不需要对场景中的运动进行精确建模,这通过基于建模的方法来展示,同时还隐式地对事件噪声进行建模。使用这个模拟器,我们使用来自大规模图像数据集的模拟数据,训练了一对关于对象检测和2D人体姿态估计的下游网络,并展示了网络推广到具有真实的事件的数据集的能力。本文中的代码和数据集可以在这里获得:https://github.com/alexzzhu/EventGAN。
Event cameras provide a number of benefits over traditional cameras, such as the ability to track incredibly fast motions, high dynamic range, and low power consumption. However, their application into computer vision problems, many of which are primarily dominated by deep learning solutions, has been limited by the lack of labeled training data for events. In this work, we propose a method which leverages the existing labeled data for images by simulating events from a pair of temporal image frames, using a convolutional neural network. We train this network on pairs of images and events, using an adversarial discriminator loss and a pair of cycle consistency losses. The cycle consistency losses utilize a pair of pre-trained self-supervised networks which perform optical flow estimation and image reconstruction from events, and constrain our network to generate events which result in accurate outputs from both of these networks. Trained fully end to end, our network learns a generative model for events from images without the need for accurate modeling of the motion in the scene, exhibited by modeling based methods, while also implicitly modeling event noise. Using this simulator, we train a pair of downstream networks on object detection and 2D human pose estimation from events, using simulated data from large scale image datasets, and demonstrate the networks' abilities to generalize to datasets with real events. The code and dataset in this paper are available here: https://github.com/alexzzhu/EventGAN.