Graph-based Spatial-temporal Feature Learning for Neuromorphic Vision Sensing

Graph-based Spatial-temporal Feature Learning for Neuromorphic Vision Sensing
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yin Bi;Aaron Chadha;Alhabib Abbas;Eirina Bourtsoulatze;Y. Andreopoulos
Yin Bi;Aaron Chadha;Alhabib Abbas;Eirina Bourtsoulatze;Y. Andreopoulos
中科院分区:
其他
文献类型:
--
作者:
Yin Bi;Aaron Chadha;Alhabib Abbas;Eirina Bourtsoulatze;Y. Andreopoulos

文献摘要

被引文献

相似文献

神经形态视觉传感(NVS)\设备将视觉信息表示为异步离散事件序列(也称为,“尖峰”)。与传统的有源像素感测(APS)不同,NVS允许显著更高的事件采样率,同时大幅提高能量效率和对照明变化的鲁棒性。然而,NVS的特征表示远远落后于其基于APS的同行,导致高级计算机视觉任务的性能较低。为了充分利用其稀疏性和异步性,我们提出了一种用于NVS的紧凑图表示,它允许使用图卷积神经网络进行端到端学习。我们将其与一个新颖的端到端特征学习框架相结合,该框架可同时适应基于外观和基于运动的任务。我们框架的核心包括一个空间特征学习模块,该模块利用残差图卷积神经网络(RG-CNN)直接从图中进行基于外观的特征的端到端学习。我们扩展了我们提出的Graph 2Grid块和时间特征学习模块,用于有效地对多个图和长时间范围的时间依赖关系进行建模。我们展示了我们的框架可以如何配置对象分类,动作识别和动作相似性标记。重要的是,我们的方法保留了尖峰事件的空间和时间相干性,同时需要更少的计算和内存。实验验证表明,我们提出的框架优于所有最近的方法在标准数据集上。最后,为了解决缺乏大型真实世界NVS数据集用于复杂识别任务的问题,我们引入、评估并提供了美国手语字母(ASL-DVS)以及人类动作数据集(UCF 101-DVS、HMDB 51-DVS和ASLAN-DVS)。
Neuromorphic vision sensing (NVS)\ devices represent visual information as sequences of asynchronous discrete events (a.k.a., "spikes") in response to changes in scene reflectance. Unlike conventional active pixel sensing (APS), NVS allows for significantly higher event sampling rates at substantially increased energy efficiency and robustness to illumination changes. However, feature representation for NVS is far behind its APS-based counterparts, resulting in lower performance in high-level computer vision tasks. To fully utilize its sparse and asynchronous nature, we propose a compact graph representation for NVS, which allows for end-to-end learning with graph convolution neural networks. We couple this with a novel end-to-end feature learning framework that accommodates both appearance-based and motion-based tasks. The core of our framework comprises a spatial feature learning module, which utilizes residual-graph convolutional neural networks (RG-CNN), for end-to-end learning of appearance-based features directly from graphs. We extend this with our proposed Graph2Grid block and temporal feature learning module for efficiently modelling temporal dependencies over multiple graphs and a long temporal extent. We show how our framework can be configured for object classification, action recognition and action similarity labeling. Importantly, our approach preserves the spatial and temporal coherence of spike events, while requiring less computation and memory. The experimental validation shows that our proposed framework outperforms all recent methods on standard datasets. Finally, to address the absence of large real-world NVS datasets for complex recognition tasks, we introduce, evaluate and make available the American Sign Language letters (ASL-DVS), as well as human action dataset (UCF101-DVS, HMDB51-DVS and ASLAN-DVS).