Spike representation of depth image sequences and its application to hand gesture recognition with spiking neural network

Spike representation of depth image sequences and its application to hand gesture recognition with spiking neural network
复制标题

DOI:
10.1007/s11760-023-02574-3
复制
发表时间:
2023-04
期刊:
Signal, Image and Video Processing
影响因子:
--
通讯作者:
Daisuke Miki;Kento Kamitsuma;Taiga Matsunaga
Daisuke Miki;Kento Kamitsuma;Taiga Matsunaga
中科院分区:
其他
文献类型:
--
作者:
Daisuke Miki;Kento Kamitsuma;Taiga Matsunaga

文献摘要

相似文献

手势在表达人们的情感和传达他们的意图方面起着重要的作用。因此,人们研究了各种方法来清楚地捕捉和理解它们。人工神经网络(ANN)由于其表达能力和易于实现而被广泛用于手势识别。然而,这项任务仍然具有挑战性,因为它需要大量的数据和计算能量。最近,低功耗的神经形态设备,使用尖峰神经网络(SNN),它可以处理时间信息,并需要较低的功耗计算,吸引了显着的研究兴趣。在这项研究中,我们提出了一种方法的尖峰表示的人类手势,并分析它们使用SNNs。SNN包括多个卷积层;当输入对应于手势的尖峰序列时,输出层中对应于每个手势的尖峰神经元激发,并且基于其激发频率对手势进行分类。利用手势的深度图像序列,研究了一种从训练图像数据生成锋电位序列的方法。手势可以通过使用代理梯度(SG)学习来训练SNN来分类。此外,通过将深度图像数据转换为尖峰序列,与ANN下的分类准确度相比,可以减少68%的训练数据量,而不会显著降低分类准确度。
Hand gestures play an important role in expressing the emotions of people and communicating their intentions. Therefore, various methods have been studied to clearly capture and understand them. Artificial neural networks (ANNs) are widely used for gesture recognition owing to their expressive power and ease of implementation. However, this task remains challenging because it requires abundant data and energy for computation. Recently, low-power neuromorphic devices that use spiking neural networks (SNNs), which can process temporal information and require lower power consumption for computing, have attracted significant research interest. In this study, we present a method for the spike representation of human hand gestures and analyzing them using SNNs. An SNN comprises multiple convolutional layers; when a sequence of spike trains corresponding to a hand gesture is inputted, the spiking neurons in the output layer corresponding to each gesture fire, and the gesture is classified based on its firing frequency. Using a sequence of depth images of hand gestures, a method to generate spike trains from the training image data was investigated. The gestures could be classified by training the SNN using surrogate gradient (SG) learning. Additionally, by converting the depth image data into spike trains, 68% of the training data volume could be reduced without significantly reducing the classification accuracy, compared to the classification accuracy under ANNs.