Surface Electromyography-Based Recognition, Synthesis, and Perception of Prosodic Subvocal Speech.

Surface Electromyography-Based Recognition, Synthesis, and Perception of Prosodic Subvocal Speech.
复制标题

DOI:
10.1044/2021_jslhr-20-00257
复制
发表时间:
2021-05
期刊:
Journal of speech, language, and hearing research : JSLHR
影响因子:
--
通讯作者:
Jennifer M. Vojtech;Michael D. Chan;Bhawna Shiwani;Serge H. Roy;J. Heaton;Geoffrey S. Meltzner;Paola Contessa;G. De Luca;R. Patel;Joshua C. Kline
Jennifer M. Vojtech;Michael D. Chan;Bhawna Shiwani;Serge H. Roy;J. Heaton;Geoffrey S. Meltzner;Paola Contessa;G. De Luca;R. Patel;Joshua C. Kline
中科院分区:
其他
文献类型:
--
作者:
Jennifer M. Vojtech;Michael D. Chan;Bhawna Shiwani;Serge H. Roy;J. Heaton;Geoffrey S. Meltzner;Paola Contessa;G. De Luca;R. Patel;Joshua C. Kline

文献摘要

相似文献

目的 本研究旨在评估一种新颖的通信系统,该系统旨在使用个性化的数字语音将来自关节肌肉的表面肌电 (sEMG) 信号转换为语音。该系统的单词识别、韵律分类和听者对合成语音的感知进行了评估。方法 进行喉部切除术的说话者 (n = 4) 和未进行喉切除术的说话者 (n = 4) 默念(默念)包含 750 个短语的语音语料库(150 个具有可变短语级别重音的短语),从面部和颈部记录 sEMG 信号。然后,通过个性化语音合成(n = 8 个合成语音)将语料库标记翻译成语音,并与每个说话者在使用其典型通信模式(n = 4 个自然语音、n = 4 个电喉 [EL] 语音)时产生的短语进行比较。初级听众 (n = 12) 在视觉排序和评分任务中评估合成语音、自然语音和 EL 语音的可接受性和可理解性,以及通过分类机制的短语重音辨别能力。结果记录的 sEMG 信号经过处理,将 sEMG 肌肉活动转化为词汇内容,并对短语级压力的变化进行分类,平均准确度分别为 96.3% (SD = 3.10%) 和 91.2% (SD = 4.46%)。合成语音的可接受性和可理解性明显高于 EL 语音,也导致了更高的短语重音分类准确性,而自然语音被评为最可接受和可理解性,具有最高的短语重音分类准确性。结论 这项概念验证研究确立了使用基于声带 sEMG 的替代交流的可行性,不仅可以用于词汇识别,还可以用于健康个体以及患有声乐障碍和残余发音功能的人的韵律交流。补充材料 https://doi.org/10.23641/asha.14558481。
Purpose This study aimed to evaluate a novel communication system designed to translate surface electromyographic (sEMG) signals from articulatory muscles into speech using a personalized, digital voice. The system was evaluated for word recognition, prosodic classification, and listener perception of synthesized speech. Method sEMG signals were recorded from the face and neck as speakers with (n = 4) and without (n = 4) laryngectomy subvocally recited (silently mouthed) a speech corpus comprising 750 phrases (150 phrases with variable phrase-level stress). Corpus tokens were then translated into speech via personalized voice synthesis (n = 8 synthetic voices) and compared against phrases produced by each speaker when using their typical mode of communication (n = 4 natural voices, n = 4 electrolaryngeal [EL] voices). Naïve listeners (n = 12) evaluated synthetic, natural, and EL speech for acceptability and intelligibility in a visual sort-and-rate task, as well as phrasal stress discriminability via a classification mechanism. Results Recorded sEMG signals were processed to translate sEMG muscle activity into lexical content and categorize variations in phrase-level stress, achieving a mean accuracy of 96.3% (SD = 3.10%) and 91.2% (SD = 4.46%), respectively. Synthetic speech was significantly higher in acceptability and intelligibility than EL speech, also leading to greater phrasal stress classification accuracy, whereas natural speech was rated as the most acceptable and intelligible, with the greatest phrasal stress classification accuracy. Conclusion This proof-of-concept study establishes the feasibility of using subvocal sEMG-based alternative communication not only for lexical recognition but also for prosodic communication in healthy individuals, as well as those living with vocal impairments and residual articulatory function. Supplemental Material https://doi.org/10.23641/asha.14558481.