Recognition of Noisy Speech: A Comparative Survey of Robust Model Architecture and Feature Enhancement

Recognition of Noisy Speech: A Comparative Survey of Robust Model Architecture and Feature Enhancement
复制标题

噪声语音识别:鲁棒模型架构和特征增强的比较调查

DOI:
--
复制
发表时间:
2009
期刊:
EURASIP Journal on Audio, Speech, and Music Processing
影响因子:
--
通讯作者:
G. Rigoll
G. Rigoll
中科院分区:
--
文献类型:
--
作者:
Björn Schuller;M. Wöllmer;T. Moosmayr;G. Rigoll

文献摘要

被引文献

相似文献

语音识别系统的性能在存在背景噪声(如汽车内的驾驶噪声)的情况下会严重下降。与现有的工作相比,我们的目标是提高噪声鲁棒性,重点放在语音识别的所有主要层次:特征提取,特征增强,语音建模和训练。因此,我们给出了一个有前途的听觉建模的概念,语音增强技术,训练策略和模型架构,这是实现在车内的数字和拼写识别任务,考虑到各种汽车类型和驾驶条件产生的噪声的概述。我们证明,联合语音和噪声建模与切换线性动态模型(SLDM)优于语音增强技术,如直方图增强(HEQ)的平均相对误差减少52.7%,在各种噪声类型和水平。将切换线性动力系统(SLDS)嵌入到切换自回归隐马尔可夫模型(SAR-HMM)中,用于加性白色高斯噪声干扰的语音。
Performance of speech recognition systems strongly degrades in the presence of background noise, like the driving noise inside a car. In contrast to existing works, we aim to improve noise robustness focusing on all major levels of speech recognition: feature extraction, feature enhancement, speech modelling, and training. Thereby, we give an overview of promising auditory modelling concepts, speech enhancement techniques, training strategies, and model architecture, which are implemented in an in-car digit and spelling recognition task considering noises produced by various car types and driving conditions. We prove that joint speech and noise modelling with a Switching Linear Dynamic Model (SLDM) outperforms speech enhancement techniques like Histogram Equalisation (HEQ) with a mean relative error reduction of 52.7% over various noise types and levels. Embedding a Switching Linear Dynamical System (SLDS) into a Switching Autoregressive Hidden Markov Model (SAR-HMM) prevails for speech disturbed by additive white Gaussian noise.