Classification of Voice Disorders Using a One-Dimensional Convolutional Neural Network

Classification of Voice Disorders Using a One-Dimensional Convolutional Neural Network
复制标题

DOI:
10.1016/j.jvoice.2020.02.009
复制
发表时间:
2022-01-04
期刊:
影响因子:
2.2
通讯作者:
Hori, Ryusuke
Hori, Ryusuke
中科院分区:
医学3区
文献类型:
--
作者:
Fujimura, Shintaro;Kojima, Tsuyoshi;Hori, Ryusuke

文献摘要

被引文献

相似文献

目标。听觉-知觉声音分析是量化病理性声音质量的标准方法,但知觉评分是基于主观评估的,因此不同的检查人员可能会有所不同。虽然许多声学指标已被研究用于病理性嗓音的客观评估,但个别病例的声学指标的解释是困难的,临床医生也没有广泛使用这项技术。目的建立一维卷积神经网络(1D-CNN)模型直接判别病理性嗓音等级、粗糙度、喘息、乏力、劳损(GRBAS)量表评分的标准化方法。我们利用1,377个元音/a/持续发音的语音样本构建了原始数据集。每个语音样本由三位专家根据GRBAS量表进行评分,并以中位数作为正确答案标签。我们设计了一个端到端的1D-CNN模型,输入的原始语音波形的帧宽度为9600个样本。使用我们的原始数据集分别对每个GRBAS类别进行训练,并采用五次交叉验证方法对模型性能进行测试。确定了测试数据集的准确度、F1分数和二次加权Cohen‘s kappa。G量表的指标显示出最平衡的模型性能,具有很高的准确性(0.771)和基本一致(kappa=0.710)。R量表模型具有较高的准确性(0.765)和F1评分(0.743),具有中等的一致性(kappa=0.536)。S量表的准确率(0.883)和F1评分(0.865)最高,而科恩卡伯量表的准确率(0.190)最低。端到端1D-CNN模型可以评估总体病理语音质量,其可靠性可与人类评估相媲美。机器学习模型训练和评估的效率与数据集质量密切相关。
Objectives. Auditory-perceptual voice analysis is a standard method for quantifying pathological voice quality, but perceptual ratings are based on subjective evaluations and therefore may vary among examiners. Although many acoustic metrics have been studied for potential use in the objective evaluation of pathological voices, the interpretation of acoustic metrics in individual cases is difficult and the technique is not widely used by clinicians. The aim of this study was to establish standardized methods to discriminate grade, roughness, breathiness, asthenia, strain (GRBAS) scale scores of pathological voices directly using one-dimensional convolutional neural network (1D-CNN) models.Methods. We constructed an original dataset utilizing 1,377 voice samples of sustained phonation of the vowel /a/. Each voice sample was rated by three experts according to the GRBAS scale and the median values were used as the correct answer label. We designed an end-to-end 1D-CNN model with a raw voice waveform input having a frame width of 9,600 samples. The models were trained with our original dataset for each GRBAS category individually and the model performance was tested by the five-fold cross validation method.Results. The accuracy, F1 score, and quadratic weighted Cohen's kappa for the testing dataset were determined. The metrics for the G scale showed the most balanced model performance, with high accuracy (0.771) and substantial agreement (kappa = 0.710). The model for the R scale had relatively high accuracy (0.765) and F1 score (0.743) with moderate agreement (kappa = 0.536). The accuracy (0.883) and the F1 score (0.865) for the S scale were the highest among the five categories, whereas the Cohen's kappa was the lowest (0.190).Conclusions. The end-to-end 1D-CNN models can evaluate overall pathological voice quality with a reliability comparable to human evaluations. The efficiency with which the machine learning models can be trained and evaluated is closely related to the dataset quality.