Neuromorphic Vision Sensing for CNN-based Action Recognition

Neuromorphic Vision Sensing for CNN-based Action Recognition
复制标题

DOI:
10.1109/icassp.2019.8683606
复制
发表时间:
2019-05
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Aaron Chadha;Yin Bi;Alhabib Abbas;Y. Andreopoulos
Aaron Chadha;Yin Bi;Alhabib Abbas;Y. Andreopoulos
中科院分区:
其他
文献类型:
--
作者:
Aaron Chadha;Yin Bi;Alhabib Abbas;Y. Andreopoulos

文献摘要

被引文献

相似文献

神经形态视觉感测(NVS)硬件现在作为一种低功耗/高速视觉感测技术获得了关注,该技术规避了传统有源像素感测(APS)相机的限制。虽然已经结合NVS研究了对象检测和跟踪模型,但目前NVS用于更高级别语义任务(如动作识别)的工作很少。与最近考虑流域(光流到运动矢量)之间的均匀传输的工作相反,我们建议将NVS仿真器嵌入到多模式迁移学习框架中,该框架执行从光流到NVS的异构传输。我们的框架的潜力通过以下事实展示:我们基于NVS的结果第一次实现了与基于运动矢量或光流的方法相当的动作识别性能(即,UCF-101上的精度在光流I3 D的8.8%以内),与标准Brox光流相比,NVS仿真器和NVS相机硬件分别提供快3到6个数量级的帧生成。除了这个显著的优势,我们的CNN处理被发现具有最低的总GFLOP计数对所有竞争的方法(高达7.7倍的复杂性节省相比,I3 D与光流)。
Neuromorphic vision sensing (NVS) hardware is now gaining traction as a low-power/high-speed visual sensing technology that circumvents the limitations of conventional active pixel sensing (APS) cameras. While object detection and tracking models have been investigated in conjunction with NVS, there is currently little work on NVS for higher-level semantic tasks, such as action recognition. Contrary to recent work that considers homogeneous transfer between flow domains (optical flow to motion vectors), we propose to embed an NVS emulator into a multi-modal transfer learning framework that carries out heterogeneous transfer from optical flow to NVS. The potential of our framework is showcased by the fact that, for the first time, our NVS-based results achieve comparable action recognition performance to motion-vector or optical-flow based methods (i.e., accuracy on UCF-101 within 8.8% of I3D with optical flow), with the NVS emulator and NVS camera hardware offering 3 to 6 orders of magnitude faster frame generation (respectively) compared to standard Brox optical flow. Beyond this significant advantage, our CNN processing is found to have the lowest total GFLOP count against all competing methods (up to 7.7 times complexity saving compared to I3D with optical flow).