Robust speech recognition using inter-speaker and intra-speaker adaptation

Robust speech recognition using inter-speaker and intra-speaker adaptation
复制标题

使用说话者间和说话者内自适应的鲁棒语音识别

DOI:
10.21437/icslp.2002-59
复制
发表时间:
2002
影响因子:
3
通讯作者:
N. Minematsu
N. Minematsu
中科院分区:
--
文献类型:
--
作者:
Baojie Li;K. Hirose;N. Minematsu

文献摘要

参考文献

被引文献

相似文献

在语音识别中,使用MLLR和MAP等说话人自适应技术可以很好地应对说话人之间的变化。然而,在处理阅读风格以外的语音时,如会话语音、情绪语音等,目前的识别系统即使经过说话人的适应,也不能达到令人满意的效果。针对这种情况,本文提出了两层自适应方法,在两个层次上应用自适应技术来处理说话人之间和说话人内部的变化。首先将扬声器独立模型适应于特定扬声器以生成扬声器依赖模型。然后,在将训练数据分成几个类别后,使用分类到每个类别的数据进一步适应说话人依赖模型(类别依赖模型)。使用说话人依赖模型和每个类别依赖模型并行识别,并选择可能性最大的结果作为最终识别结果。对不同情绪的语音(输入语音的情绪未知)进行了识别实验,结果表明该方法优于传统的基于mllr的说话人自适应方法。
Inter-speaker variation can be coped rather well in speech recognition by speaker adaptation techniques such as MLLR and MAP. However, when dealing with speech other than reading style, such as conversational speech, emotional speech and so on, current recognition systems cannot achieve a satisfactory performance even after speaker adaptation. In view of this situation, two-level adaptation method was newly proposed, where adaptation technique was applied in two levels to handle inter-speaker and in-tra-speaker variations. A speaker independent model is first adapted to a specific speaker to generate a speaker dependent model. Then, after classifying the training data into several categories, the speaker dependent model is further adapted to each category using data classified to it (category dependent model). The recognition is done in parallel using the speaker dependent model and each category dependent model, and the result with highest likelihood is selected as the final recognition result. Recognition experiments were conducted for speech with various emotions (emotion of input speech is unknown), and the results showed that the proposed method outperformed the conventional MLLR-based speaker adaptation.
DOI: 10.1006/csla.1995.0010
发表时间: 1995-04-01
影响因子: 4.3
作者:
LEGGETTER, CJ;WOODLAND, PC
通讯作者: WOODLAND, PC