A 3DCNN-LSTM Multi-Class Temporal Segmentation for Hand Gesture Recognition

A 3DCNN-LSTM Multi-Class Temporal Segmentation for Hand Gesture Recognition
复制标题

DOI:
10.3390/electronics11152427
复制
发表时间:
2022-08-01
期刊:
影响因子:
2.9
通讯作者:
Bharath, Anil A.
Bharath, Anil A.
中科院分区:
工程技术3区
文献类型:
--
作者:
Gionfrida, Letizia;Rusli, Wan M. R.;Bharath, Anil A.

文献摘要

被引文献

相似文献

本文介绍了一种多类手势识别模型,识别一组手势序列从二维RGB视频记录,使用连续帧的外观和时空参数。该分类器利用基于卷积的网络与长短期记忆单元相结合。为了利用对大规模数据集的需求,该模型在公共数据集上部署训练,采用一种称为迁移学习的技术来微调相关手势的架构。在64个批量上进行的验证曲线表明,22名参与者的准确度为93.95%(+/- 0.37),平均Jaccard指数为0.812(+/- 0.105)。微调的架构说明了用一小部分数据(113,410个完全标记的图像帧)改进模型以覆盖以前未知的手势的可能性。这项工作的主要贡献包括一个由单目RGB视频序列驱动的自定义手势识别网络,该网络优于以前的时间分割模型,采用了一个小型架构,便于广泛采用。
This paper introduces a multi-class hand gesture recognition model developed to identify a set of hand gesture sequences from two-dimensional RGB video recordings, using both the appearance and spatiotemporal parameters of consecutive frames. The classifier utilizes a convolutional-based network combined with a long-short-term memory unit. To leverage the need for a large-scale dataset, the model deploys training on a public dataset, adopting a technique known as transfer learning to fine-tune the architecture on the hand gestures of relevance. Validation curves performed over a batch size of 64 indicate an accuracy of 93.95% (+/- 0.37) with a mean Jaccard index of 0.812 (+/- 0.105) for 22 participants. The fine-tuned architecture illustrates the possibility of refining a model with a small set of data (113,410 fully labelled image frames) to cover previously unknown hand gestures. The main contribution of this work includes a custom hand gesture recognition network driven by monocular RGB video sequences that outperform previous temporal segmentation models, embracing a small-sized architecture that facilitates wide adoption.