Tackling Speaking Mode Varieties in EMG-Based Speech Recognition

Tackling Speaking Mode Varieties in EMG-Based Speech Recognition
复制标题

DOI:
10.1109/tbme.2014.2319000
复制
发表时间:
2014-10-01
影响因子:
4.6
通讯作者:
Schultz, Tanja
Schultz, Tanja
中科院分区:
工程技术2区
文献类型:
--
作者:
Wand, Michael;Janke, Matthias;Schultz, Tanja

文献摘要

被引文献

相似文献

肌电图(EMG)无声语音识别器是一种通过捕捉人类发音肌肉的电位来识别语音的系统,从而使用户能够无声地交流。在建立了基于基线肌电图的连续语音识别器之后,本文研究了说话模式的变化,即可听和无声语音之间的差异,这些差异会降低识别的准确性。我们引入了多模式系统,允许在可听语音和无声语音之间无缝切换,研究了量化说话模式差异的不同措施,并提出了频谱映射算法,该算法将无声语音的单词错误率(WER)相对提高了14.3%。无声语音的最佳平均WER为34.7%,有声语音的最佳平均WER为16.8%。
An electromyographic (EMG) silent speech recognizer is a system that recognizes speech by capturing the electric potentials of the human articulatory muscles, thus enabling the user to communicate silently. After having established a baseline EMG-based continuous speech recognizer, in this paper, we investigate speaking mode variations, i.e., discrepancies between audible and silent speech that deteriorate recognition accuracy. We introduce multimode systems that allow seamless switching between audible and silent speech, investigate different measures which quantify speaking mode differences, and present the spectral mapping algorithm, which improves the word error rate (WER) on silent speech by up to 14.3% relative. Our best average silent speech WER is 34.7%, and our best WER on audibly spoken speech is 16.8%.