GSVD-based optimal filtering for single and multimicrophone speech enhancement

GSVD-based optimal filtering for single and multimicrophone speech enhancement
复制标题

DOI:
10.1109/tsp.2002.801937
复制
发表时间:
2002-09
期刊:
IEEE Trans. Signal Process.
影响因子:
--
通讯作者:
S. Doclo;M. Moonen
S. Doclo;M. Moonen
中科院分区:
其他
文献类型:
--
作者:
S. Doclo;M. Moonen

文献摘要

被引文献

相似文献

提出了一种基于广义奇异值分解(GSVD)的多麦克风语音信号增强算法。这种基于GSVD的多麦克风算法可以被认为是用于增强噪声语音信号的单麦克风信号子空间算法的扩展,并且当无法观察到期望的响应信号时,相当于特定的最优滤波问题。最佳滤波器可以被写为语音和噪声数据矩阵的广义奇异向量和奇异值的函数。一些对称性的单麦克风和多麦克风的最佳滤波器,这是有效的白色噪声的情况下,以及有色噪声的情况下,得出。此外,一些单麦克风信号子空间算法的平均步骤进行了检查,导致的结论是,这种平均操作是不必要的,甚至次优。对于简单的情况下,我们考虑本地化的源和没有多径传播,基于GSVD的最佳滤波技术表现出波束形成器的空间方向性图案。当比较针对现实情况的降噪性能时,仿真表明,基于GSVD的最优滤波技术对于所有混响时间具有比标准固定和自适应波束成形技术更好的性能,并且其对于与标称情况的偏差更鲁棒,例如,在未校准的麦克风阵列中遇到。
A generalized singular value decomposition (GSVD) based algorithm is proposed for enhancing multimicrophone speech signals degraded by additive colored noise. This GSVD-based multimicrophone algorithm can be considered to be an extension of the single-microphone signal subspace algorithms for enhancing noisy speech signals and amounts to a specific optimal filtering problem when the desired response signal cannot be observed. The optimal filter can be written as a function of the generalized singular vectors and singular values of a speech and noise data matrix. A number of symmetry properties are derived for the single-microphone and multimicrophone optimal filter, which are valid for the white noise case as well as for the colored noise case. In addition, the averaging step of some single-microphone signal subspace algorithms is examined, leading to the conclusion that this averaging operation is unnecessary and even suboptimal. For simple situations, where we consider localized sources and no multipath propagation, the GSVD-based optimal filtering technique exhibits the spatial directivity pattern of a beamformer. When comparing the noise reduction performance for realistic situations, simulations show that the GSVD-based optimal filtering technique has a better performance than standard fixed and adaptive beamforming techniques for all reverberation times and that it is more robust to deviations from the nominal situation, as, e.g., encountered in uncalibrated microphone arrays.