An AI-based Approach for Improved Sign Language Recognition using Multiple Videos

An AI-based Approach for Improved Sign Language Recognition using Multiple Videos
复制标题

DOI:
10.1007/s11042-021-11830-y
复制
发表时间:
2022-02-28
影响因子:
3.6
通讯作者:
Clark, Addison
Clark, Addison
中科院分区:
计算机科学4区
文献类型:
--
作者:
Dignan, Cameron;Perez, Eliud;Clark, Addison

文献摘要

被引文献

相似文献

有听力和言语障碍的人在沟通方面面临重大障碍。手语知识可以帮助减轻这些障碍,但大多数健全人,包括亲戚、朋友和护理人员,无法理解手语。自动化工具的可用性可以让残疾人及其周围的人在各种情况下与非签名者进行无处不在的沟通。目前有两种主要的手语手势识别方法。第一种是基于硬件的方法,涉及手套或其他硬件来跟踪手部位置并确定手势。第二种是基于软件的方法,其中拍摄手部视频并使用计算机视觉技术对手势进行分类。然而,某些硬件(例如手机的内部传感器或佩戴在手臂上用于跟踪肌肉数据的设备)不太准确,而且佩戴它们可能会很麻烦或不舒服。另一方面,基于软件的方法取决于照明条件以及手与背景之间的对比度。我们提出了一种混合方法,利用低成本传感硬件并将其与智能符号识别算法相结合,目标是开发更高效的手势识别系统。使用支持向量机方法的基于 Myo 带的方法的准确度仅为 49%。基于软件的方法使用卷积神经网络(CNN)和循环神经网络(RNN)方法来训练基于Myo的模块,并在我们的实验中实现了超过80%的准确率。我们的方法结合了这两种方法并显示了改进的潜力。我们的实验是使用由多个视频生成的九个手势的数据集完成的,每个手势重复五次,对基于软件和基于硬件的模块总共进行了 45 次试验。除了显示每种方法的性能之外,我们的结果还表明,通过进一步改进的硬件模块,可以显着提高组合方法的准确性。
People with hearing and speaking disabilities face significant hurdles in communication. The knowledge of sign language can help mitigate these hurdles, but most people without disabilities, including relatives, friends, and care providers, cannot understand sign language. The availability of automated tools can allow people with disabilities and those around them to communicate ubiquitously and in a variety of situations with non-signers. There are currently two main approaches for recognizing sign language gestures. The first is a hardware-based approach, involving gloves or other hardware to track hand position and determine gestures. The second is a software-based approach, where a video is taken of the hands and gestures are classified using computer vision techniques. However, some hardware, such as a phone's internal sensor or a device worn on the arm to track muscle data, is less accurate, and wearing them can be cumbersome or uncomfortable. The software-based approach, on the other hand, is dependent on the lighting conditions and on the contrast between the hands and the background. We propose a hybrid approach that takes advantage of low-cost sensory hardware and combines it with a smart sign-recognition algorithm with the goal of developing a more efficient gesture recognition system. The Myo band-based approach using the Support Vector Machine method achieves an accuracy of only 49%. The software-based approach uses the Convolutional Neural Network (CNN) and Recurrent Neural Network (RNN) methods to train the Myo-based module and achieves an accuracy of over 80% in our experiments. Our method combines the two approaches and shows the potential for improvement. Our experiments are done with a dataset of nine gestures generated from multiple videos, each repeated five times for a total of 45 trials for both the software-based and hardware-based modules. Apart from showing the performance of each approach, our results show that with a more improved hardware module, the accuracy of the combined approach can be significantly improved.