Monaural Speech Enhancement Based on Spectrogram Decomposition for Convolutional Neural Network-sensitive Feature Extraction

Monaural Speech Enhancement Based on Spectrogram Decomposition for Convolutional Neural Network-sensitive Feature Extraction
复制标题

DOI:
10.21437/interspeech.2022-11268
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Hao Shi;Longbiao Wang;Sheng Li;J. Dang;Tatsuya Kawahara
Hao Shi;Longbiao Wang;Sheng Li;J. Dang;Tatsuya Kawahara
中科院分区:
其他
文献类型:
--
作者:
Hao Shi;Longbiao Wang;Sheng Li;J. Dang;Tatsuya Kawahara

文献摘要

相似文献

许多最先进的语音增强(SE)系统最近使用卷积神经网络(CNN)来提取多尺度特征图。然而,CNN更多地依赖于局部纹理而不是全局形状,这更容易受到降级的频谱图的影响,并且可能无法捕获语音的详细结构。虽然一些两级系统将第一级增强的和原始的噪声频谱图同时馈送到第二级,但这不能保证对第二级的足够指导,因为第一级频谱图不能提供精确的频谱细节。为了使CNN能够感知清晰的语音分量边界信息,我们根据第一阶段的掩码值将特征图与包含明显语音分量的频谱图组合在一起。对应于大于特定阈值的掩模的位置被提取为特征图。这些特征图通过忽略其他特征使语音分量的边界信息变得明显,从而使CNN对输入特征敏感。在VB数据集上的实验表明,通过适当的分解次数,该方法可以提高SE性能,可以提供0.15的PESQ改善。此外,该方法在光谱细节恢复方面更为有效。
Many state-of-the-art speech enhancement (SE) systems have recently used convolutional neural networks (CNNs) to extract multi-scale feature maps. However, CNN relies more on local texture than global shape, which is more susceptible to degraded spectrogram and may fail to capture the detailed structure of speech. Although some two-stage systems feed the first-stage enhanced and original noisy spectrograms to the second stage simultaneously, this does not guarantee sufficient guidance for the second stage since the first-stage spectrogram can not provide precise spectral details. In order to allow CNNs to perceive clear speech component boundary information, we compose feature maps with spectrograms containing evident speech components according to the mask value from the first stage. The positions corresponding to the mask greater than certain thresholds are extracted as feature maps. These feature maps make the boundary information of speech components obvious by ignoring others, thus making CNNs sensitive to input features. Experiments on the VB dataset show that with a proper decomposition numbers, the proposed method can enhance SE performance, which can provide 0.15 PESQ improvement. Besides, the proposed method is more effective for spectral detail recovery.