Noise‐robust speech recognition using multi‐band spectral features

Noise‐robust speech recognition using multi‐band spectral features
复制标题

使用多频带频谱特征的抗噪声语音识别

DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
S. Furui
S. Furui
中科院分区:
--
文献类型:
--
作者:
Y. Nishimura;T. Shinozaki;K. Iwano;S. Furui

文献摘要

被引文献

相似文献

在大多数最先进的自动语音识别(ASR)系统中,语音被转换为MFCC(Mel频率倒谱系数)向量的时间函数。然而,使用MFCC的问题是,即使噪声被限制在窄频带内,噪声效应也会扩散到所有系数上。如果直接使用频谱特征,则可以避免这样的问题,并且因此可以预期增加对噪声的鲁棒性。虽然已经进行了各种研究,使用谱域特征,识别性能的改善已被报道仅在有限的噪声条件下。本文提出了一种新的多波段ASR方法,使用一种新的对数谱域特征。为了提高鲁棒性,对数谱特征通过应用三个过程进行归一化:减去每个帧的平均对数能量,强调谱峰,以及减去在话语上平均的对数谱平均值。光谱分量似然值...
In most of the state‐of‐the‐art automatic speech recognition (ASR) systems, speech is converted into a time function of the MFCC (Mel Frequency Cepstrum Coefficient) vector. However, the problem with using the MFCC is that noise effects spread over all the coefficients even when the noise is limited within a narrow frequency band. If a spectrum feature is directly used, such a problem can be avoided and thus robustness against noise could be expected to increase. Although various researches on using spectral domain features have been conducted, improvement of recognition performances has been reported only in limited noise conditions. This paper proposes a novel multi‐band ASR method using a new log‐spectral domain feature. In order to increase the robustness, log‐spectrum features are normalized by applying the three processes: subtracting the mean log‐energy for each frame, emphasizing spectral peaks, and subtracting the log‐spectral mean averaged over an utterance. Spectral component likelihood values ...