Near real-time ASL recognition using a millimeter wave radar

Near real-time ASL recognition using a millimeter wave radar
复制标题

DOI:
10.1117/12.2588616
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
Oladipupo O. Adeoluwa;Sean J. Kearney;Emre Kurtoğlu;Charles Connors;S. Gurbuz
Oladipupo O. Adeoluwa;Sean J. Kearney;Emre Kurtoğlu;Charles Connors;S. Gurbuz
中科院分区:
其他
文献类型:
--
作者:
Oladipupo O. Adeoluwa;Sean J. Kearney;Emre Kurtoğlu;Charles Connors;S. Gurbuz

文献摘要

相似文献

针对聋人群体的技术研究大多集中在使用视频或可穿戴设备进行翻译上。据报道,传感器增强型手套比基于摄像头的系统具有更高的手势识别率;然而,它们无法捕捉通过头部和身体运动所表达的信息。手套还具有侵入性,会妨碍用户的日常生活,而摄像头可能会引发隐私问题,且在黑暗中无效。相比之下,射频传感器是非接触式、非侵入性的,即使被黑客攻击也不会泄露隐私信息。尽管射频传感器无法测量面部表情或手部形状(完整翻译需要这些信息),但本文旨在利用射频传感器进行近乎实时的美国手语(ASL)识别,以设计智能聋人空间。通过这种方式,我们希望聋人群体能够从技术进步中受益,从而切实提高他们的生活质量。更具体地说,本文研究了机器学习和深度学习架构的近乎实时实现,以用于连续的美国手语手势识别。我们使用一个60GHz的射频传感器,它发射调频连续波(FMWC波形)。射频传感器能够获取光学或可穿戴设备无法获取的独特信息源:即通过微多普勒特征对运动的运动学模式进行可视化呈现。微多普勒是指围绕中心多普勒频移出现的频率调制,它是由偏离主平移运动的旋转或振动运动引起的。在先前的工作中,我们表明从射频数据计算出的分形复杂度可用于区分手语和日常活动,并且射频数据能够揭示语言特性,如协同发音。我们还表明,机器学习可用于以99%的准确率区分本土聋人美国手语使用者的手语和听力正常个体的模仿手语(或仿手语)。因此,模仿手语数据对于直接训练深度模型是无效的。但是,对抗学习可用于将模仿手语转换为类似于本土手语,或者,物理感知生成模型可用于合成美国手语微多普勒特征以训练深度神经网络。通过这些方法,我们对20个美国手语手势实现了超过90%的识别准确率。然而,在自然环境中,需要分类算法的近乎实时实现,以及以连续和顺序的方式处理数据流的能力。在这项工作中,我们专注于对先前工作朝着这个目标进行扩展,并比较在树莓派或Jetson板等平台上嵌入深度神经网络(DNNs)的各种方法的效果。我们研究了用于优化嵌入式微多普勒分析的深度神经网络的大小和计算复杂度的方法、网络压缩方法以及它们所产生的连续美国手语识别性能。
Most research in technologies for the Deaf community have focused on translation using either video or wearable devices. Sensor-augmented gloves have been reported to yield higher gesture recognition rates than camera-based systems; however, they cannot capture information expressed through head and body movement. Gloves are also intrusive and inhibit users in their pursuit of normal daily life, while cameras can raise concerns over privacy and are ineffective in the dark. In contrast, RF sensors are non-contact, non-invasive and do not reveal private information even if hacked. Although RF sensors are unable to measure facial expressions or hand shapes, which would be required for complete translation, this paper aims to exploit near real-time ASL recognition using RF sensors for the design of smart Deaf spaces. In this way, we hope to enable the Deaf community to benefit from advances in technologies that could generate tangible improvements in their quality of life. More specifically, this paper investigates near real-time implementation of machine learning and deep learning architectures for the purpose of sequential ASL signing recognition. We utilize a 60 GHz RF sensor which transmits a frequency modulation continuous wave (FMWC waveform). RF sensors can acquire a unique source of information that is inaccessible to optical or wearable devices: namely, a visual representation of the kinematic patterns of motion via the micro-Doppler signature. Micro-Doppler refers to frequency modulations that appear about the central Doppler shift, which are caused by rotational or vibrational motions that deviate from principle translational motion. In prior work, we showed that fractal complexity computed from RF data could be used to discriminate signing from daily activities and that RF data could reveal linguistic properties, such as coarticulation. We have also shown that machine learning can be used to discriminate with 99% accuracy the signing of native Deaf ASL users from that of copysigning (or imitation signing) by hearing individuals. Therefore, imitation signing data is not effective for directly training deep models. But, adversarial learning can be used to transform imitation signing to resemble native signing, or, alternatively, physics-aware generative models can be used to synthesize ASL micro-Doppler signatures for training deep neural networks. With such approaches, we have achieved over 90% recognition accuracy of 20 ASL signs. In natural environments, however, near real-time implementations of classification algorithms are required, as well as an ability to process data streams in a continuous and sequential fashion. In this work, we focus on extensions of our prior work towards this aim, and compare the efficacy of various approaches for embedding deep neural networks (DNNs) on platforms such as a Raspberry Pi or Jetson board. We examine methods for optimizing the size and computational complexity of DNNs for embedded micro-Doppler analysis, methods for network compression, and their resulting sequential ASL recognition performance.