Development of sEMG sensors and algorithms for silent speech recognition.

Development of sEMG sensors and algorithms for silent speech recognition.
复制标题

DOI:
10.1088/1741-2552/aac965
复制
发表时间:
2018-08
影响因子:
4
通讯作者:
Kline JC
Kline JC
中科院分区:
工程技术2区
文献类型:
--
作者:
Meltzner GS;Heaton JT;Deng Y;De Luca G;Roy SH;Kline JC

文献摘要

参考文献

被引文献

相似文献

语音是人类最自然的交流形式之一,因此通过自动语音识别(ASR)为人机交互提供了一种有吸引力的模式。然而,ASR的局限性,包括在环境噪声的存在下,有限的隐私和穷人的可访问性显着的言语障碍,激励了替代的非声学模态的无声或无声的语音识别(SSR)的需要。我们开发了一种新的面部和颈部佩戴传感器和信号处理算法系统,能够完全从面部和颈部肌肉记录的表面肌电(sEMG)信号中识别无声的单词和短语。这些算法是通过不断发展的语音识别模型进行战略性开发的:首先通过从sEMG信号中提取语音相关特征来识别孤立的单词,然后使用语法模型从sEMG信号的模式中识别单词序列,最后使用基于音素的模型识别以前未训练的单词的词汇表。最终的识别算法与专门设计的多点微型传感器集成,这些传感器可以以灵活的几何形状排列,以记录面部和颈部小咬合肌的高保真sEMG信号测量。我们在一系列无声语音实验中测试了传感器和算法系统,这些实验涉及从2200个单词的词汇表中生成的1200多个短语,并实现了8.9%的单词错误率(91.1%的识别率),远远超过了该领域以前的尝试。这些结果表明,我们的系统的可行性作为一种替代的通信模式的多种应用程序,包括:人与语音障碍后喉切除术;军事人员需要免提隐蔽通信;或消费者在需要的隐私,而在公共场合发言的移动的电话。
Speech is among the most natural forms of human communication, thereby offering an attractive modality for human–machine interaction through automatic speech recognition (ASR). However, the limitations of ASR—including degradation in the presence of ambient noise, limited privacy and poor accessibility for those with significant speech disorders—have motivated the need for alternative non-acoustic modalities of subvocal or silent speech recognition (SSR). We have developed a new system of face- and neck-worn sensors and signal processing algorithms that are capable of recognizing silently mouthed words and phrases entirely from the surface electromyographic (sEMG) signals recorded from muscles of the face and neck that are involved in the production of speech. The algorithms were strategically developed by evolving speech recognition models: first for recognizing isolated words by extracting speech-related features from sEMG signals, then for recognizing sequences of words from patterns of sEMG signals using grammar models, and finally for recognizing a vocabulary of previously untrained words using phoneme-based models. The final recognition algorithms were integrated with specially designed multi-point, miniaturized sensors that can be arranged in flexible geometries to record high-fidelity sEMG signal measurements from small articulator muscles of the face and neck. We tested the system of sensors and algorithms during a series of subvocal speech experiments involving more than 1200 phrases generated from a 2200-word vocabulary and achieved an 8.9%-word error rate (91.1% recognition rate), far surpassing previous attempts in the field. These results demonstrate the viability of our system as an alternative modality of communication for a multitude of applications including: persons with speech impairments following a laryngectomy; military personnel requiring hands-free covert communication; or the consumer in need of privacy while speaking on a mobile phone in public.
DOI: 10.1007/bf02345373
发表时间: 2001-07-01
影响因子: 3.2
作者:
Chan, ADC;Englehart, K;Lovely, DF
通讯作者: Lovely, DF
DOI: 10.1007/s11517-007-0168-z
发表时间: 2007-05-01
影响因子: 3.2
作者:
Roy, S. H.;De Luca, G.;De Luca, C. J.
通讯作者: De Luca, C. J.
DOI: 10.1109/taslp.2017.2752365
发表时间: 2017-12-01
影响因子: 5.4
作者:
Schultz, Tanja;Wand, Michael;Brumberg, Jonathan S.
通讯作者: Brumberg, Jonathan S.
DOI: 10.1016/j.csl.2010.06.003
发表时间: 2011-04-01
影响因子: 4.3
作者:
Povey, Daniel;Burget, Lukas;Thomas, Samuel
通讯作者: Thomas, Samuel
DOI: 10.1016/j.specom.2009.12.002
发表时间: 2010-04-01
影响因子: 3.2
作者:
Schultz, Tanja;Wand, Michael
通讯作者: Wand, Michael