Modeling coarticulation in EMG-based continuous speech recognition

Modeling coarticulation in EMG-based continuous speech recognition
复制标题

DOI:
10.1016/j.specom.2009.12.002
复制
发表时间:
2010-04-01
影响因子:
3.2
通讯作者:
Wand, Michael
Wand, Michael
中科院分区:
计算机科学3区
文献类型:
--
作者:
Schultz, Tanja;Wand, Michael

文献摘要

被引文献

相似文献

本文讨论了表面肌电信号在自动语音识别中的应用。在面部肌肉处捕获的肌电信号记录人类发音器官的活动,并且因此允许追溯语音信号,即使其是无声地说出的。由于语音在传播之前被捕获,因此产生的信号不会被环境噪声掩盖。由此产生的无声语音界面有可能克服传统语音驱动界面的主要局限性:它不容易受到任何环境噪声的影响,允许无声地传输机密信息,并且不会打扰旁观者。我们描述了我们的新方法,用于在基于EMG的语音识别中建模协同发音的语音特征捆绑,并在EMG-PIT语料库上报告结果,一个多说话人的大词汇量数据库的无声和可听EMG语音录音,我们最近收集。我们的研究结果依赖于说话者和独立于说话者的设置表明,语音特征的相互依存关系的建模减少了超过33%的相对基准系统的单词错误率。我们的最终系统实现了10%的单词错误率为最好的公认的扬声器上的101个单词的词汇任务,使基于EMG的语音识别在一个有用的范围内的无声语音接口的应用。(C)2009爱思唯尔有限公司版权所有。
This paper discusses the use of surface electromyography for automatic speech recognition. Electromyographic signals captured at the facial muscles record the activity of the human articulatory apparatus and thus allow to trace back a speech signal even if it is spoken silently. Since speech is captured before it gets airborne, the resulting signal is not masked by ambient noise. The resulting Silent Speech Interface has the potential to overcome major limitations of conventional speech-driven interfaces: it is not prone to any environmental noise, allows to silently transmit confidential information, and does not disturb bystanders.We describe our new approach of phonetic feature bundling for modeling coarticulation in EMG-based speech recognition and report results on the EMG-PIT corpus, a multiple speaker large vocabulary database of silent and audible EMG speech recordings, which we recently collected. Our results on speaker-dependent and speaker-independent setups show that modeling the interdependence of phonetic features reduces the word error rate of the baseline system by over 33% relative. Our final system achieves 10% word error rate for the best-recognized speaker on a 101-word vocabulary task, bringing EMG-based speech recognition within a useful range for the application of Silent Speech Interfaces. (C) 2009 Elsevier B.V. All rights reserved.