User-Independent American Sign Language Alphabet Recognition Based on Depth Image and PCANet Features

User-Independent American Sign Language Alphabet Recognition Based on Depth Image and PCANet Features
复制标题

DOI:
10.1109/access.2019.2938829
复制
发表时间:
2019-01-01
期刊:
影响因子:
3.9
通讯作者:
Almotairi, Sultan
Almotairi, Sultan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Aly, Walaa;Aly, Saleh;Almotairi, Sultan

文献摘要

被引文献

相似文献

手语是聋人与正常人交流最自然、最有效的方式。美国手语字母识别(即手指拼写)使用无标记视觉传感器是一个具有挑战性的任务,由于困难的手分割和外观之间的差异签名。现有的基于颜色的手语识别系统面临着复杂背景、手部分割、类内和类间差异大等挑战。在本文中,我们提出了一个新的用户独立的识别系统,美国手语字母使用深度图像捕获的低成本微软Kinect深度传感器。利用深度信息而不是彩色图像克服了许多问题,因为它们对照明和背景变化的鲁棒性。通过对深度图像应用简单的预处理算法可以分割手部区域。应用使用卷积神经网络架构的特征学习,而不是经典的手工特征提取方法。使用简单的无监督主成分分析网络(PCANet)深度学习架构有效地学习从分割的手提取的局部特征。提出了两种学习PCANet模型的策略,即从所有用户的样本中训练单个PCANet模型和为每个用户分别训练单独的PCANet模型。然后使用线性支持向量机(SVM)分类器识别提取的特征。所提出的方法的性能进行评估,使用公共数据集的真实的深度图像从不同的用户。实验结果表明,该方法的性能优于最先进的识别精度采用留一法评价策略。
Sign language is the most natural and effective way for communications among deaf and normal people. American Sign Language (ASL) alphabet recognition (i.e. fingerspelling) using marker-less vision sensor is a challenging task due to the difficulties in hand segmentation and appearance variations among signers. Existing color-based sign language recognition systems suffer from many challenges such as complex background, hand segmentation, large inter-class and intra-class variations. In this paper, we propose a new user independent recognition system for American sign language alphabet using depth images captured from the low-cost Microsoft Kinect depth sensor. Exploiting depth information instead of color images overcomes many problems due to their robustness against illumination and background variations. Hand region can be segmented by applying a simple preprocessing algorithm over depth image. Feature learning using convolutional neural network architectures is applied instead of the classical hand-crafted feature extraction methods. Local features extracted from the segmented hand are effectively learned using a simple unsupervised Principal Component Analysis Network (PCANet) deep learning architecture. Two strategies of learning the PCANet model are proposed, namely to train a single PCANet model from samples of all users and to train a separate PCANet model for each user, respectively. The extracted features are then recognized using linear Support Vector Machine (SVM) classifier. The performance of the proposed method is evaluated using public dataset of real depth images captured from various users. Experimental results show that the performance of the proposed method outperforms state-of-the-art recognition accuracy using leave-one-out evaluation strategy.