Decoding silent speech commands from articulatory movements through soft magnetic skin and machine learning

Decoding silent speech commands from articulatory movements through soft magnetic skin and machine learning
复制标题

通过软磁皮肤和机器学习从发音运动解码无声语音命令

DOI:
10.1039/d3mh01062g
复制
发表时间:
2023
期刊:
影响因子:
13.3
通讯作者:
Yao, Shanshan
Yao, Shanshan
中科院分区:
材料科学1区
文献类型:
--
作者:
Dong, Penghao;Li, Yizong;Chen, Si;Grafstein, Justin T.;Khan, Irfaan;Yao, Shanshan

文献摘要

相似文献

当基于声学的语音通信不可靠、不合适或不受欢迎时,人们追求无声语音接口来恢复具有语音障碍的个人的语音通信,并促进直观通信。然而,当前的无声语音方法面临着一些挑战,包括体积庞大、引人注目、准确率低、可移植性有限以及对干扰的敏感性。在这项工作中,我们提出了一个无线的,不显眼的,健壮的无声语音接口,用于跟踪和解码与语音相关的颞颌关节运动。我们的解决方案采用了放置在耳朵后面的单一软磁性皮肤,用于无线和社会可接受的静默语音识别。开发的系统缓解了与基于面部佩戴的传感器的现有接口相关的几个问题,包括大量传感器、面部高度可见的接口以及传感器和数据采集组件之间的突出互连。利用基于机器学习的信号处理技术,获得了良好的语音识别准确率(音素准确率为93.2%,相同视位组单词列表的准确率为87.3%)。此外,报告的无声语音接口显示出对来自环境环境和用户日常运动的噪声的稳健性。最后,通过两个演示-支持无声语音的智能手机助手和支持无声语音的无人机控制-展示了它在辅助技术和人机交互方面的潜力。
Silent speech interfaces have been pursued to restore spoken communication for individuals with voice disorders and to facilitate intuitive communications when acoustic-based speech communication is unreliable, inappropriate, or undesired. However, the current methodology for silent speech faces several challenges, including bulkiness, obtrusiveness, low accuracy, limited portability, and susceptibility to interferences. In this work, we present a wireless, unobtrusive, and robust silent speech interface for tracking and decoding speech-relevant movements of the temporomandibular joint. Our solution employs a single soft magnetic skin placed behind the ear for wireless and socially acceptable silent speech recognition. The developed system alleviates several concerns associated with existing interfaces based on face-worn sensors, including a large number of sensors, highly visible interfaces on the face, and obtrusive interconnections between sensors and data acquisition components. With machine learning-based signal processing techniques, good speech recognition accuracy is achieved (93.2% accuracy for phonemes, and 87.3% for a list of words from the same viseme groups). Moreover, the reported silent speech interface demonstrates robustness against noises from both ambient environments and users’ daily motions. Finally, its potential in assistive technology and human–machine interactions is illustrated through two demonstrations – silent speech enabled smartphone assistants and silent speech enabled drone control.