Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition

Two-Step Acoustic Model Adaptation for Dysarthric Speech Recognition
复制标题

DOI:
10.1109/icassp40776.2020.9053725
复制
发表时间:
2020-05
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
R. Takashima;T. Takiguchi;Y. Ariki
R. Takashima;T. Takiguchi;Y. Ariki
中科院分区:
其他
文献类型:
--
作者:
R. Takashima;T. Takiguchi;Y. Ariki

文献摘要

相似文献

本文介绍了一种用于依赖于说话人的构音障碍语音识别系统的模型自适应方法。我们在本文中关注的构音障碍是由手足徐动型脑瘫引起的,这种疾病会导致患者出现不自主的肌肉运动。因此,构音障碍者的语音往往不稳定,传统的自动语音识别(ASR)系统难以识别。模型适应方法是一种可能的解决方案,该方法使 ASR 模型适应构音障碍语音。然而,由于构音障碍者和非构音障碍者之间的说话风格差异如此显着,传统的适应方法无法使模型充分适应构音障碍的语音。在我们提出的两步模型适应方法中,ASR 模型首先适应多个构音障碍说话者的一般说话风格,然后将适应的模型进一步适应目标说话者。从我们对 ASR 任务的实验来看,我们的两步适应方法比传统的一步适应方法表现出更好的性能。
This paper introduces a model adaptation approach for a speaker-dependent dysarthric speech recognition system. The dysarthria we focus on in this paper is caused by athetoid cerebral palsy, which causes involuntary muscle movements in those with the disease. For this reason, the dysarthric people’s speech is often unstable and difficult for conventional automatic speech recognition (ASR) systems to recognize. A model-adaptation approach, which adapts an ASR model to dysarthric speech, is one possible solution. However, because the difference in speaking styles between dysarthric and non-dysarthric people is so significant, the conventional adaptation method is not able to sufficiently adapt the model to the dysarthric speech. In our proposed two-step model-adaptation approach, an ASR model is first adapted to the general speaking style of multiple dysarthric speakers, and then the adapted model is further adapted for the target speaker. From our experiments on an ASR task, our two-step adaptation approach showed better performance than a conventional one-step adaptation approach.