Bayesian Feature Enhancement for Reverberation and Noise Robust Speech Recognition

Bayesian Feature Enhancement for Reverberation and Noise Robust Speech Recognition
复制标题

用于混响和噪声鲁棒语音识别的贝叶斯特征增强

DOI:
--
复制
发表时间:
2013
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Reinhold Häb
Reinhold Häb
中科院分区:
--
文献类型:
--
作者:
Volker Leutnant;A. Krueger;Reinhold Häb

文献摘要

被引文献

相似文献

在这一贡献中,我们将先前提出的用于增强混响对数Mel功率谱系数的贝叶斯方法扩展到额外的背景噪声补偿。采用了最近提出的一种观测模型,该模型的时变观测误差统计量是干净语音特征向量的后验概率密度函数推断的副积。此外,通过使用观测模型的递归公式来实现计算工作量和存储器需求的减少。首先在人工产生噪声混响数据的连接数字识别任务中对所提出的算法的性能进行了实验研究。结果表明,在低信噪比条件下,与时不变观测误差模型相比,采用时变观测误差模型可以显著降低误码率。进一步的实验是在混响和噪声环境中记录的5000字任务上进行的。实验结果表明,该方法在实际数据上具有较好的字错误率。
In this contribution we extend a previously proposed Bayesian approach for the enhancement of reverberant logarithmic mel power spectral coefficients for robust automatic speech recognition to the additional compensation of background noise. A recently proposed observation model is employed whose time-variant observation error statistics are obtained as a side product of the inference of the a posteriori probability density function of the clean speech feature vectors. Further a reduction of the computational effort and the memory requirements are achieved by using a recursive formulation of the observation model. The performance of the proposed algorithms is first experimentally studied on a connected digits recognition task with artificially created noisy reverberant data. It is shown that the use of the time-variant observation error model leads to a significant error rate reduction at low signal-to-noise ratios compared to a time-invariant model. Further experiments were conducted on a 5000 word task recorded in a reverberant and noisy environment. A significant word error rate reduction was obtained demonstrating the effectiveness of the approach on real-world data.