Differentiable Consistency Constraints for Improved Deep Speech Enhancement

Differentiable Consistency Constraints for Improved Deep Speech Enhancement
复制标题

用于改进深度语音增强的可微一致性约束

DOI:
10.1109/icassp.2019.8682783
复制
发表时间:
2018
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
R. Saurous
R. Saurous
中科院分区:
--
文献类型:
--
作者:
Scott Wisdom;J. Hershey;K. Wilson;J. Thorpe;Michael Chinen;Brian Patton;R. Saurous

文献摘要

被引文献

相似文献

近年来,深度网络通过将语音增强视为数据驱动的模式识别问题,使语音增强得到了显著改善。在许多现代增强系统中,大量数据用于训练深度网络来估计复值短时傅立叶变换(STFT)的掩码,以抑制噪声并保留语音。然而,目前的掩蔽方法往往忽略了两个重要的约束:STFT一致性和混合一致性。在没有STFT一致性的情况下,系统的输出不一定是时域信号的STFT,并且在没有混合一致性的情况下,估计的源的总和不一定等于输入混合。此外,之前唯一应用混合一致性的方法使用实值掩码;对于复值掩码,混合一致性已被忽略。
In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transforms (STFTs) to suppress noise and preserve speech. However, current masking approaches often neglect two important constraints: STFT consistency and mixture consistency. Without STFT consistency, the system’s output is not necessarily the STFT of a time-domain signal, and without mixture consistency, the sum of the estimated sources does not necessarily equal the input mixture. Furthermore, the only previous approaches that apply mixture consistency use real-valued masks; mixture consistency has been ignored for complex-valued masks.