Estimators of The Magnitude-Squared Spectrum and Methods for Incorporating SNR Uncertainty.

Estimators of The Magnitude-Squared Spectrum and Methods for Incorporating SNR Uncertainty.
复制标题

DOI:
10.1109/tasl.2010.2082531
复制
发表时间:
2011-07-01
期刊:
IEEE transactions on audio, speech, and language processing
影响因子:
--
通讯作者:
Loizou PC
Loizou PC
中科院分区:
其他
文献类型:
--
作者:
Lu Y;Loizou PC

文献摘要

被引文献

相似文献

基于噪声语音信号的幅度平方谱可计算为(纯净)信号幅度平方谱与噪声幅度平方谱之和这一假设,推导出幅度平方谱的统计估计量。基于高斯统计模型推导出最大后验(MAP)和最小均方误差(MMSE)估计量。发现MAP估计量的增益函数与在计算听觉场景分析(CASA)中广泛使用的理想二值掩蔽(IdBM)中所使用的增益函数相同。因此,它是二值的,如果局部信噪比超过0 dB,则取值为1,否则取值为0。通过将局部瞬时信噪比建模为F分布随机变量,推导出包含信噪比不确定性的软掩蔽方法。特别是,根据局部信噪比超过0 dB的先验概率对噪声幅度平方谱进行加权的软掩蔽方法被证明与维纳增益函数相同。结果表明,就产生更低的残留噪声和更低的语音失真而言,所提出的估计量比传统的MMSE谱功率估计量能产生明显更好的语音质量。
Statistical estimators of the magnitude-squared spectrum are derived based on the assumption that the magnitude-squared spectrum of the noisy speech signal can be computed as the sum of the (clean) signal and noise magnitude-squared spectra. Maximum a posterior (MAP) and minimum mean square error (MMSE) estimators are derived based on a Gaussian statistical model. The gain function of the MAP estimator was found to be identical to the gain function used in the ideal binary mask (IdBM) that is widely used in computational auditory scene analysis (CASA). As such, it was binary and assumed the value of 1 if the local SNR exceeded 0 dB, and assumed the value of 0 otherwise. By modeling the local instantaneous SNR as an F-distributed random variable, soft masking methods were derived incorporating SNR uncertainty. The soft masking method, in particular, which weighted the noisy magnitude-squared spectrum by the a priori probability that the local SNR exceeds 0 dB was shown to be identical to the Wiener gain function. Results indicated that the proposed estimators yielded significantly better speech quality than the conventional MMSE spectral power estimators, in terms of yielding lower residual noise and lower speech distortion.