Vocal frequency estimation and voicing state prediction with surface EMG pattern recognition

Vocal frequency estimation and voicing state prediction with surface EMG pattern recognition
复制标题

DOI:
10.1016/j.specom.2014.04.004
复制
发表时间:
2014-09-01
影响因子:
3.2
通讯作者:
Chau, Tom
Chau, Tom
中科院分区:
计算机科学3区
文献类型:
--
作者:
De Armas, Winston;Mamun, Khondaker A.;Chau, Tom

文献摘要

被引文献

相似文献

大多数喉切除者在全喉切除术后使用电子喉作为主要的言语交流方式。然而,典型的电子喉具有单调的音调和需要手动控制的不便。本文提出了潜在的模式识别,以支持电子喉的使用预测基本频率(F0)和发声状态(VS)从表面肌电图的舌骨下和舌骨上肌肉,以及从呼吸痕迹。在这项研究中,从舌骨下和舌骨上肌肉群和呼吸痕迹的表面肌电信号收集了10个健全的成年男性(18- 60岁)。受试者完成了三种语音任务:声调、连读和短语。从肌电信号和呼吸道信号中提取特征,采用径向基函数核的支持向量机分类器预测F0和发声状态。对于90-360 Hz范围内的声音频率的估计,平均均方根误差为2.81 +/-0.6。平均交叉验证(CV)准确度为78.05 +/- 6.3%,从EMG和65.24 +/- 7.8%,从呼吸道预测发声状态。与依赖肌内电极(侵入性)的研究相比,所提出的方法具有非侵入性的优点,同时仍保持高于机会的准确性。颈部肌肉表面肌电信号的模式分类在预测发声过程中的基频和发声状态方面具有优势,鼓励进一步研究电子喉和无声语音接口的自动音高调制。(C)2014爱思唯尔有限公司版权所有。
The majority of laryngectomees use the electrolarynx as their primary mode of verbal communication after total laryngectomy surgery. However, the archetypal electrolarynx suffers from a monotonous tone and the inconvenience of requiring manual control. This paper presents the potential of pattern recognition to support electrolarynx use by predicting fundamental frequency (F0) and voicing state (VS) from surface EMG of the infrahyoid and suprahyoid muscles, as well as from a respiratory trace. In this study, surface EMG signals from the infrahyoid and suprahyoid muscle groups and respiratory trace were collected from 10 able-bodied, adult males (18- 60 years old). Participants performed three kinds of vocal tasks tones, legatos and phrases. Signal features were extracted from the EMG and respiratory trace, and a Support Vector Machine (SVM) classifier with radial basis function kernels was employed to predict F0 and voicing state. An average root mean squared error of 2.81 +/- 0.6 semitones was achieved for the estimation of vocal frequency in the range of 90-360 Hz. An average cross-validation (CV) accuracy of 78.05 +/- 6.3% was achieved for the prediction of voicing state from EMG and 65.24 +/- 7.8% from the respiratory trace. The proposed method has the advantage of being non-invasive compared with studies that relied on intramuscular electrodes (invasive), while still maintaining an accuracy above chance. Pattern classification of neck-muscle surface EMG has merit in the prediction of fundamental frequency and voicing state during vocalization, encouraging further study of automatic pitch modulation for electrolarynges and silent speech interfaces. (C) 2014 Elsevier B.V. All rights reserved.