Rapid environment adaptation method based on HMM composition with prior noise GMM and multi‐SNR models for noisy speech recognition

Rapid environment adaptation method based on HMM composition with prior noise GMM and multi‐SNR models for noisy speech recognition
复制标题

基于 HMM 与先验噪声 GMM 和多 SNR 模型组合的快速环境适应方法,用于噪声语音识别

DOI:
10.1002/ecjb.20093
复制
发表时间:
2004
期刊:
Electronics and Communications in Japan Part Ii-electronics
影响因子:
--
通讯作者:
Satoshi Nakamura
Satoshi Nakamura
中科院分区:
--
文献类型:
--
作者:
M. Ida;Satoshi Nakamura

文献摘要

参考文献

被引文献

相似文献

语音识别系统在真实环境中使用时,不可避免地会在输入语音中存在周围环境噪声,从而降低识别性能。在大多数情况下,很难预测噪声的混合情况,输入信号与声学模型之间的噪声环境差异是导致识别性能下降的原因之一。因此,需要建立一个对各种噪声混合具有鲁棒性的声学模型。噪声混合问题可以分为两个方面,即噪声种类的多样化和信噪比的多样化。本文将基于加权自适应噪声GMM的HMM组合应用于第一个问题,将多信噪比路径模型应用于第二个问题。在嘈杂环境下的语音识别实验中,使用旅行会话任务和AURORA2任务对这两种方法的组合进行了性能评估。当在AURORA2任务中使用1秒的自适应数据,信噪比为5 dB时,识别率比基线HMM提高了53%。这与传统HMM合成中使用10秒适应数据的情况相对应。©2004 Wiley期刊公司电子工程学报,2004,31 (6):393 - 398;在线发表于Wiley InterScience (www.interscience.wiley.com)。DOI 10.1002 / ecjb.20093
In the use of speech recognition systems in a real environment, it is inevitable that surrounding environmental noise is present in the input speech, which degrades recognition performance. It is difficult in most cases to predict the mixing of the noise, and the discrepancy of noise environments between the input signal and the acoustic model is a reason for degradation of recognition performance. Consequently, it is desirable to construct an acoustic model which is robust to the mixing of various kinds of noise. The problem of noise mixture can be divided into two aspects, namely, diversified kinds of noise and diversified values of the SNR. In this paper, HMM composition using weight adaptation of the noise GMM is applied to the first problem, and the multi-SNR path model is applied to the second problem. Performance evaluation is performed for a combination of these two approaches in a speech recognition experiment in a noisy environment, using the travel conversation task and the AURORA2 task. When 1 second of adaptation data is used in the AURORA2 task for SNR = 5 dB, the recognition rate is improved by 53% compared to the baseline HMM. This corresponds to the case in which 10 seconds of adaptation data is used in conventional HMM composition. © 2004 Wiley Periodicals, Inc. Electron Comm Jpn Pt 2, 87(6): 39–48, 2004; Published online in Wiley InterScience (www.interscience.wiley.com). DOI 10.1002/ecjb.20093
DOI: 10.1006/csla.1995.0010
发表时间: 1995-04-01
影响因子: 4.3
作者:
LEGGETTER, CJ;WOODLAND, PC
通讯作者: WOODLAND, PC