Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach

Detection of Pathological Voice Using Cepstrum Vectors: A Deep Learning Approach
复制标题

DOI:
10.1016/j.jvoice.2018.02.003
复制
发表时间:
2019-09-01
期刊:
影响因子:
2.2
通讯作者:
Wang, Chi-Te
Wang, Chi-Te
中科院分区:
医学3区
文献类型:
--
作者:
Fang, Shih-Hau;Tsao, Yu;Wang, Chi-Te

文献摘要

被引文献

相似文献

目标.嗓音疾病的计算机检测已引起了学术界和临床界的极大兴趣,人们希望在内窥镜确诊之前提供一种有效的嗓音疾病筛查方法。本研究提出了一种基于深度学习的病态语音检测方法,并与其他自动分类算法进行了比较。本研究回顾性地收集了一所三级教学医院嗓音门诊的60个正常嗓音样本和402个临床上常见的8种嗓音疾病的病理嗓音样本。我们从持续元音的3秒样本中提取Mel频率倒谱系数。三种机器学习算法,即深度神经网络(DNN),支持向量机和高斯混合模型的性能进行了评估的基础上,五重交叉验证。使用来自MEEI(马萨诸塞州眼耳医院)嗓音障碍数据库的集体病例来验证分类机制的性能。实验结果表明,DNN优于高斯混合模型和支持向量机。基于三个具有代表性的Mel倒谱系数特征,该方法对男性和女性嗓音病变的检测准确率分别达到94.26%和90.52%。在MEEI数据库上进行验证时,DNN的分类准确率(99.32%)也高于其他两种分类算法。通过堆叠多层具有优化权重的神经元,所提出的DNN算法可以充分利用声学特征并有效区分正常和病态语音样本。基于这项初步研究,未来的研究可能会从实验室和临床角度探索DNN的更多应用。
Objectives. Computerized detection of voice disorders has attracted considerable academic and clinical interest in the hope of providing an effective screening method for voice diseases before endoscopic confirmation. This study proposes a deep-learning-based approach to detect pathological voice and examines its performance and utility compared with other automatic classification algorithms.Methods. This study retrospectively collected 60 normal voice samples and 402 pathological voice samples of 8 common clinical voice disorders in a voice clinic of a tertiary teaching hospital. We extracted Mel frequency cepstral coefficients from 3-second samples of a sustained vowel. The performances of three machine learning algorithms, namely, deep neural network (DNN), support vector machine, andGaussian mixture model, were evaluated based on a fivefold cross-validation. Collective cases from the voice disorder database of MEEI (Massachusetts Eye and Ear Infirmary) were used to verify the performance of the classification mechanisms.Results. The experimental results demonstrated that DNN outperforms Gaussian mixture model and support vector machine. Its accuracy in detecting voice pathologies reached 94.26% and 90.52% in male and female subjects, based on three representative Mel frequency cepstral coefficient features. When applied to the MEEI database for validation, theDNN also achieved a higher accuracy (99.32%) than the other two classification algorithms.Conclusions. By stacking several layers of neurons with optimized weights, the proposed DNN algorithm can fully utilize the acoustic features and efficiently differentiate between normal and pathological voice samples. Based on this pilot study, future researchmay proceed to explore more application ofDNNfrom laboratory and clinical perspectives.