Computing MMSE Estimates and Residual Uncertainty Directly in the Feature Domain of ASR using STFT Domain Speech Distortion Models

Computing MMSE Estimates and Residual Uncertainty Directly in the Feature Domain of ASR using STFT Domain Speech Distortion Models
复制标题

使用 STFT 域语音失真模型直接在 ASR 特征域中计算 MMSE 估计和残余不确定性

DOI:
10.1109/tasl.2013.2244085
复制
发表时间:
2013
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
R. Orglmeister
R. Orglmeister
中科院分区:
--
文献类型:
--
作者:
Ramón Fernández Astudillo;R. Orglmeister

文献摘要

被引文献

相似文献

在本文中,我们演示了如何不确定性传播允许计算的最小均方误差(MMSE)估计在特征域的各种特征提取方法,使用短时傅立叶变换(STFT)域失真模型。除此之外,还获得了估计可靠性的衡量标准,允许特征重新估计或自动语音识别(ASR)模型的动态补偿。所提出的方法通过使用STFT不确定性传播公式的特征提取将与维纳滤波器相关联的后验分布变换。它还表明,在STFT域中的非线性估计,如Ephraim-Malah滤波器可以被看作是特殊情况下的维纳后验的传播。该方法是通过开发两个MMSE梅尔频率倒谱系数(MFCC)估计,并将它们与观测不确定性技术相结合。我们讨论了与其他MMSE-MFCC估计的相似性,并展示了所提出的方法如何在AURORA 4鲁棒ASR任务的STFT域中优于传统的MMSE估计。
In this paper we demonstrate how uncertainty propagation allows the computation of minimum mean square error (MMSE) estimates in the feature domain for various feature extraction methods using short-time Fourier transform (STFT) domain distortion models. In addition to this, a measure of estimate reliability is also attained which allows either feature re-estimation or the dynamic compensation of automatic speech recognition (ASR) models. The proposed method transforms the posterior distribution associated to a Wiener filter through the feature extraction using the STFT Uncertainty Propagation formulas. It is also shown that non-linear estimators in the STFT domain like the Ephraim-Malah filters can be seen as special cases of a propagation of the Wiener posterior. The method is illustrated by developing two MMSE-Mel-frequency Cepstral Coefficient (MFCC) estimators and combining them with observation uncertainty techniques. We discuss similarities with other MMSE-MFCC estimators and show how the proposed approach outperforms conventional MMSE estimators in the STFT domain on the AURORA4 robust ASR task.