DNN-Based Low-Musical-Noise Single-Channel Speech Enhancement Based on Higher-Order-Moments Matching

DNN-Based Low-Musical-Noise Single-Channel Speech Enhancement Based on Higher-Order-Moments Matching
复制标题

DOI:
10.1587/transinf.2021edp7041
复制
发表时间:
2021-11
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Satoshi Mizoguchi;Yuki Saito;Shinnosuke Takamichi;H. Saruwatari
Satoshi Mizoguchi;Yuki Saito;Shinnosuke Takamichi;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Satoshi Mizoguchi;Yuki Saito;Shinnosuke Takamichi;H. Saruwatari

文献摘要

相似文献

我们提出了基于深度神经网络(DNN)的语音增强,可以减少音乐噪声并获得更好的听觉印象。音乐噪声是由非线性信号处理产生的伪影,对听觉印象产生负面影响。我们的目标是开发音乐噪声的语音增强方法,抑制音乐噪声的产生,并产生感知舒适的增强语音。使用软掩模的基于DNN的语音增强实现了高降噪,但在非语音区域中产生音乐噪声。因此,首先,我们定义了基于DNN的低音乐噪声语音增强的峰度匹配。峰度是四阶矩,已知与音乐噪声的量相关。峰度匹配是DNN训练的惩罚项,用于减少音乐噪声量。我们进一步扩展该计划的矩匹配。扩展方案涉及使用阶数高于峰度的矩,并推广了传统的基于峰度匹配的无音乐噪声方法。我们制定了高阶矩匹配,并探讨如何有效地减少音乐噪声量。实验结果表明:1)峰度匹配可以在不影响噪声抑制的情况下降低音乐噪声; 2)最新发现,六阶矩匹配在峰度匹配的基础上还可以实现低音乐噪声语音增强。
SUMMARY We propose deep neural network (DNN)-based speech enhancement that reduces musical noise and achieves better auditory impressions. The musical noise is an artifact generated by nonlinear signal processing and negatively a ff ects the auditory impressions. We aim to develop musical-noise-free speech enhancement methods that suppress the musical noise generation and produce perceptually-comfortable enhanced speech. DNN-based speech enhancement using a soft mask achieves high noise reduction but generates musical noise in non-speech regions. Therefore, first, we define kurtosis matching for DNN-based low-musical-noise speech enhancement. Kurtosis is the fourth-order moment and is known to correlate with the amount of musical noise. The kurtosis matching is a penalty term of the DNN training and works to reduce the amount of musical noise. We further extend this scheme to standardized-moment matching. The extended scheme involves using moments whose orders are higher than kurtosis and generalizes the conventional musical-noise-free method based on kurtosis matching. We formulate standardized-moment matching and explore how e ff ectively the higher-order moments reduce the amount of musical noise. Experimental evaluation results 1) demonstrate that kurtosis matching can reduce musical noise without negatively a ff ecting noise suppression and 2) newly reveal that the sixth-moment matching also achieves low-musical-noise speech enhancement as well as kurtosis matching.