Mapping from Frame-Driven to Frame-Free Event-Driven Vision Systems by Low-Rate Rate Coding and Coincidence Processing-Application to Feedforward ConvNets

Mapping from Frame-Driven to Frame-Free Event-Driven Vision Systems by Low-Rate Rate Coding and Coincidence Processing-Application to Feedforward ConvNets
复制标题

DOI:
10.1109/tpami.2013.71
复制
发表时间:
2013-11-01
影响因子:
23.6
通讯作者:
Linares-Barranco, Bernabe
Linares-Barranco, Bernabe
中科院分区:
计算机科学1区
文献类型:
--
作者:
Antonio Perez-Carrasco, Jose;Zhao, Bo;Linares-Barranco, Bernabe

文献摘要

被引文献

相似文献

事件驱动的视觉传感器已经引起了许多不同研究团体的兴趣。它们提供视觉信息的方式与传统的视频系统完全不同,传统的视频系统由以给定的“帧速率”呈现的静止图像序列组成。“事件驱动的视觉传感器从生物学中获得灵感。每个像素在感觉到有意义的事情发生时发出一个事件(尖峰),而没有任何帧的概念。一种特殊类型的事件驱动传感器是所谓的动态视觉传感器(DVS),其中每个像素计算光线的相对变化或“时间对比度”。“传感器输出由代表场景中移动物体的连续像素事件流组成。像素事件相对于“现实”具有微秒延迟。这些事件可以由事件(卷积)处理器级联处理。因此,输入和输出事件流实际上在时间上是一致的,并且只要传感器提供足够有意义的事件,就可以识别对象。在本文中,我们提出了一种方法,从一个适当训练的神经网络在传统的框架驱动的表示映射到事件驱动的表示。该方法通过研究事件驱动的卷积神经网络(ConvNet)来说明,该网络被训练用于识别旋转的人体轮廓或高速扑克牌符号。事件驱动的ConvNet由从真实的DVS相机获得的记录提供。事件驱动的ConvNet由专用的事件驱动模拟器模拟,由许多事件驱动的处理模块组成,其特性从单独制造的硬件模块中获得。
Event-driven visual sensors have attracted interest from a number of different research communities. They provide visual information in quite a different way from conventional video systems consisting of sequences of still images rendered at a given "frame rate." Event-driven vision sensors take inspiration from biology. Each pixel sends out an event (spike) when it senses something meaningful is happening, without any notion of a frame. A special type of event-driven sensor is the so-called dynamic vision sensor (DVS) where each pixel computes relative changes of light or "temporal contrast." The sensor output consists of a continuous flow of pixel events that represent the moving objects in the scene. Pixel events become available with microsecond delays with respect to "reality." These events can be processed "as they flow" by a cascade of event (convolution) processors. As a result, input and output event flows are practically coincident in time, and objects can be recognized as soon as the sensor provides enough meaningful events. In this paper, we present a methodology for mapping from a properly trained neural network in a conventional frame-driven representation to an event-driven representation. The method is illustrated by studying event-driven convolutional neural networks (ConvNet) trained to recognize rotating human silhouettes or high speed poker card symbols. The event-driven ConvNet is fed with recordings obtained from a real DVS camera. The event-driven ConvNet is simulated with a dedicated event-driven simulator and consists of a number of event-driven processing modules, the characteristics of which are obtained from individually manufactured hardware modules.