Robust feature extraction for continuous speech recognition using the MVDR spectrum estimation method

Robust feature extraction for continuous speech recognition using the MVDR spectrum estimation method
复制标题

DOI:
10.1109/tasl.2006.876776
复制
发表时间:
2007-01-01
影响因子:
--
通讯作者:
Rao, Bhaskar D.
Rao, Bhaskar D.
中科院分区:
其他
文献类型:
--
作者:
Dharanipragada, Satya;Yapanel, Umit H.;Rao, Bhaskar D.

文献摘要

被引文献

相似文献

本文描述了一种用于连续语音识别的鲁棒特征提取技术。该技术的核心是频谱估计的最小方差无失真响应 (MVDR) 方法。我们考虑以两种方式合并感知信息:1)在计算 MVDR 功率谱之后,2)直接在 MVDR 频谱估计期间。我们表明,将感知信息直接纳入频谱估计可以显着提高鲁棒性和计算效率。我们使用 Fisher 线性判别测量分析了特征的类可分离性和说话人可变性特性,并表明这些特征比广泛使用的梅尔频率倒谱系数 (MFCC) 特征提供了更好的类可分离性和更好的说话人相关信息抑制。我们在四个不同的任务上评估该技术:车载语音识别任务、Aurora-2 匹配任务、华尔街日报 (WSJ) 任务和总机任务。在大多数情况下,新的特征提取技术比 MFCC 和感知线性预测 (PLP) 特征提取技术具有更低的字错误率。统计显着性检验表明,这种改进在高噪声条件下最为显着。因此,该技术提高了对噪声的鲁棒性,而不会牺牲清洁条件下的性能。
This paper describes a robust feature extraction technique for continuous speech recognition. Central to the technique is the minimum variance distortionless response (MVDR) method of spectrum estimation. We consider incorporating perceptual information in two ways: 1) after the MVDR power spectrum is computed and 2) directly during the MVDR spectrum estimation. We show that incorporating perceptual information directly into the spectrum estimation improves both robustness and computational efficiency significantly. We analyze the class separability and speaker variability properties of the features using a Fisher linear discriminant measure and show that these features provide better class separability and better suppression of speaker-dependent information than the widely used mel frequency cepstral coefficient (MFCC) features. We evaluate the technique on four different tasks: an in-car speech recognition task, the Aurora-2 matched task, the Wall Street Journal (WSJ) task, and the Switchboard task. The new feature extraction technique gives lower word-error-rates than the MFCC and perceptual linear prediction (PLP) feature extraction techniques in most cases. Statistical significance tests reveal that the improvement is most significant in high noise conditions. The technique thus provides improved robustness to noise without sacrificing performance in clean conditions.