Semi-Supervised Joint Enhancement of Spectral and Cepstral Sequences of Noisy Speech

Semi-Supervised Joint Enhancement of Spectral and Cepstral Sequences of Noisy Speech
复制标题

DOI:
10.21437/interspeech.2016-1286
复制
发表时间:
2016-08
期刊:
--
影响因子:
--
通讯作者:
Li Li-Li;H. Kameoka;T. Higuchi;H. Saruwatari
Li Li-Li;H. Kameoka;T. Higuchi;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Li Li-Li;H. Kameoka;T. Higuchi;H. Saruwatari

文献摘要

相似文献

虽然使用非负矩阵分解(NMF)的频谱域语音增强算法在信号恢复精度(例如,信噪比)方面是强大的,但它们并不一定会导致特征域中增强语音质量的改善。这意味着,天真地使用这些算法作为前端处理,例如语音识别和语音转换并不总是导致令人满意的结果。为了解决这一问题,本文提出了一种新的方法,通过优化一个组合目标函数,该目标函数由在谱域定义的基于nmf的模型拟合准则和在倒谱域定义的基于高斯混合模型(GMM)的概率分布组成,旨在共同增强含噪语音的频谱序列和倒谱序列。我们为该目标函数导出了一个新的极大量器,这使得我们可以推导出一个基于极大量最小化方案的保证收敛的迭代优化算法。实验结果表明,该方法在信失真比和倒谱距离方面都优于传统的NMF方法。
While spectral domain speech enhancement algorithms using non-negative matrix factorization (NMF) are powerful in terms of signal recovery accuracy (e.g., signal-to-noise ratio), they do not necessarily lead to an improvement in the quality of the enhanced speech in the feature domain. This implies that naively using these algorithms as front-end processing for e.g., speech recognition and speech conversion does not always lead to satisfactory results. To address this problem, this paper proposes a novel method that aims to jointly enhance the spectral and cepstral sequences of noisy speech, by optimizing a combined objective function consisting of an NMF-based model-fitting criterion defined in the spectral domain and a Gaussian mixture model (GMM)-based probability distribution defined in the cepstral domain. We derive a novel majorizer for this objective function, which allows us to derive a convergence-guaranteed iterative algorithm based on a majorization-minimization scheme for the optimization. Experimental results revealed that the proposed method outperformed the conventional NMF approach in terms of both signalto-distortion ratio and cepstral distance.