Single-channel speech separation with memory-enhanced recurrent neural networks

Single-channel speech separation with memory-enhanced recurrent neural networks
复制标题

DOI:
10.1109/icassp.2014.6854294
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
F. Weninger;F. Eyben;Björn Schuller
F. Weninger;F. Eyben;Björn Schuller
中科院分区:
其他
文献类型:
--
作者:
F. Weninger;F. Eyben;Björn Schuller

文献摘要

被引文献

相似文献

在本文中,我们提出了使用长短期记忆递归神经网络的语音增强。训练网络以从噪声语音特征预测干净语音以及噪声特征,并且从这些特征构造幅度域软掩模。广泛的测试运行在73 k噪声和混响的话语从视听兴趣语料库的自发的,情绪化的彩色语音,退化的几个小时的真实的噪声录音,包括固定和非固定源和卷积噪声从亚琛房间脉冲响应数据库。在结果中,所提出的方法被示出为在低信噪比下提供上级降噪,同时在较高信噪比下产生非常小的伪影,从而在源失真比方面大幅优于无监督幅度域谱减法。
In this paper we propose the use of Long Short-Term Memory recurrent neural networks for speech enhancement. Networks are trained to predict clean speech as well as noise features from noisy speech features, and a magnitude domain soft mask is constructed from these features. Extensive tests are run on 73 k noisy and reverberated utterances from the Audio-Visual Interest Corpus of spontaneous, emotionally colored speech, degraded by several hours of real noise recordings comprising stationary and non-stationary sources and convolutive noise from the Aachen Room Impulse Response database. In the result, the proposed method is shown to provide superior noise reduction at low signal-to-noise ratios while creating very little artifacts at higher signal-to-noise ratios, thereby outperforming unsupervised magnitude domain spectral subtraction by a large margin in terms of source-distortion ratio.