Improving Acoustic Models in TORGO Dysarthric Speech Database

Improving Acoustic Models in TORGO Dysarthric Speech Database
复制标题

DOI:
10.1109/tnsre.2018.2802914
复制
发表时间:
2018-03-01
影响因子:
4.9
通讯作者:
Umesh, S.
Umesh, S.
中科院分区:
工程技术2区
文献类型:
--
作者:
Joy, Neethu Mariam;Umesh, S.

文献摘要

被引文献

相似文献

辅助性语言技术可以改善构音障碍患者的生活质量。构音障碍是一种运动语言障碍。在本文中,我们探索了多种方法来改进高斯混合模型和基于深度神经网络(DNN)的隐马尔可夫模型(HMM)自动语音识别系统。与之前在TORGO中构建此类系统的尝试相比,这项工作显示了显著的改进。我们通过调整不同的声学模型参数来训练特定于说话人的声学模型,使用说话人归一化倒谱特征,并使用dropout和序列辨别策略构建复杂的DNN-HMM模型。使用广义蒸馏框架,利用来自困难语音的特定信息,对来自困难语音和正常语音的音频文件训练的DNN模型进一步改进了重度和重度中度困难说话者的DNN- hmm模型。据我们所知,本文给出了迄今为止TORGO数据库的最佳识别精度。
Assistive speech-based technologies can improve the quality of life for people affected with dysarthria, a motor speech disorder. In this paper, we explore multiple ways to improve Gaussian mixture model and deep neural network (DNN) based hidden Markov model (HMM) automatic speech recognition systems for TORGO dysarthric speech database. This work shows significant improvements over the previous attempts in building such systems in TORGO. We trained speaker-specific acoustic models by tuning various acoustic model parameters, using speaker normalized cepstral features and building complex DNN-HMM models with dropout and sequence-discrimination strategies. The DNN-HMM models for severe and severe-moderate dysarthric speakers were further improved by leveraging specific information from dysarthric speech to DNN models trained on audio files from both dysarthric and normal speech, using generalized distillation framework. To the best of our knowledge, this paper presents the best recognition accuracies for TORGO database till date.