Silent Speech Recognition as an Alternative Communication Device for Persons with Laryngectomy.

Silent Speech Recognition as an Alternative Communication Device for Persons with Laryngectomy.
复制标题

DOI:
10.1109/taslp.2017.2740000
复制
发表时间:
2017-12
期刊:
IEEE/ACM transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Kline JC
Kline JC
中科院分区:
其他
文献类型:
--
作者:
Meltzner GS;Heaton JT;Deng Y;De Luca G;Roy SH;Kline JC

文献摘要

被引文献

相似文献

每年有数以千计的人因为创伤或疾病而需要手术切除他们的喉咙(发声盒),因此需要替代的语音源或辅助设备来进行口头交流。虽然喉切除术后失去了自然发声,但大多数控制语音清晰度的肌肉仍然完好无损。可以从颈部和面部记录语音肌肉的表面肌电活动,并将其用于自动语音识别,提供语音到文本或合成语音作为另一种交流手段。这是正确的,即使当说话是嘴或以无声(低声)的方式,使其成为一个适当的沟通平台后,喉切除。在这项研究中,8个人在全喉切除后至少6个月被记录在他们的面部(4个)和颈部(4个)上的8个sEMG传感器上,同时阅读从2500个单词的词汇中构建的短语。一组独特的短语被用来为英语中39个常用音素中的每一个训练基于音素的识别模型,其余的短语被用于测试基于从正在运行的语音中的音素识别的模型的单词识别。在完整的8个传感器集合中,单词错误率平均为10.3%(前4名参与者的平均错误率为9.5%),当传感器集合减少到每个人4个位置时(n=7),错误率为13.6%。本研究为基于表面肌电信号的咽部语音识别提供了一个令人信服的概念验证,具有进一步提高识别性能的强大潜力。
Each year thousands of individuals require surgical removal of their larynx (voice box) due to trauma or disease, and thereby require an alternative voice source or assistive device to verbally communicate. Although natural voice is lost after laryngectomy, most muscles controlling speech articulation remain intact. Surface electromyographic (sEMG) activity of speech musculature can be recorded from the neck and face, and used for automatic speech recognition to provide speech-to-text or synthesized speech as an alternative means of communication. This is true even when speech is mouthed or spoken in a silent (subvocal) manner, making it an appropriate communication platform after laryngectomy. In this study, 8 individuals at least 6 months after total laryngectomy were recorded using 8 sEMG sensors on their face (4) and neck (4) while reading phrases constructed from a 2,500-word vocabulary. A unique set of phrases were used for training phoneme-based recognition models for each of the 39 commonly used phonemes in English, and the remaining phrases were used for testing word recognition of the models based on phoneme identification from running speech. Word error rates were on average 10.3% for the full 8-sensor set (averaging 9.5% for the top 4 participants), and 13.6% when reducing the sensor set to 4 locations per individual (n=7). This study provides a compelling proof-of-concept for sEMG-based alaryngeal speech recognition, with the strong potential to further improve recognition performance.