A Signal Subspace Speech Enhancement Approach Based on Joint Low-Rank and Sparse Matrix Decomposition

A Signal Subspace Speech Enhancement Approach Based on Joint Low-Rank and Sparse Matrix Decomposition
复制标题

DOI:
10.1515/aoa-2016-0024
复制
发表时间:
2016-01
影响因子:
0.9
通讯作者:
Chengli Sun;Jianxiao Xie;Y. Leng
Chengli Sun;Jianxiao Xie;Y. Leng
中科院分区:
物理与天体物理4区
文献类型:
--
作者:
Chengli Sun;Jianxiao Xie;Y. Leng

文献摘要

被引文献

相似文献

基于子空间的方法已被有效地用于从含噪语音样本中估计增强语音。在传统的子空间方法中,关键的一步是通过子空间分解将与信号和噪声相关的两个不变子空间分裂,这通常是通过奇异值分解或特征值分解来实现的。然而,这些分解算法对大的损坏的存在高度敏感,导致在低信噪比(SNR)的情况下增强的语音中存在大量的残余噪声。提出了一种基于联合低阶稀疏矩阵分解(JLSMD)的语音增强方法。在该方法中,我们首先将受污染的数据构造为Toeplitz矩阵,并估计其对底层干净语音矩阵的有效秩值。然后利用JLSMD对子空间进行分解,其中分解后的低阶部分对应于增强语音,稀疏部分对应于噪声信号。对于高斯白噪声和真实世界的噪声,都进行了大量的实验。实验结果表明,在多种强噪声环境下,该方法比传统方法具有更低的残馀噪声和更低的语音失真。
Subspace-based methods have been effectively used to estimate enhanced speech from noisy speech samples. In the traditional subspace approaches, a critical step is splitting of two invariant subspaces associated with signal and noise via subspace decomposition, which is often performed by singular-value decomposition or eigenvalue decomposition. However, these decomposition algorithms are highly sensitive to the presence of large corruptions, resulting in a large amount of residual noise within enhanced speech in low signal-to-noise ratio (SNR) situations. In this paper, a joint low-rank and sparse matrix decomposition (JLSMD) based subspace method is proposed for speech enhancement. In the proposed method, we firstly structure the corrupted data as a Toeplitz matrix and estimate its effective rank value for the underlying clean speech matrix. Then the subspace decomposition is performed by means of JLSMD, where the decomposed low-rank part corresponds to enhanced speech and the sparse part corresponds to noise signal, respectively. An extensive set of experiments have been carried out for both of white Gaussian noise and real-world noise. Experimental results show that the proposed method performs better than conventional methods in many types of strong noise conditions, in terms of yielding less residual noise and lower speech distortion.