Differentiable Consistency Constraints for Improved Deep Speech Enhancement
Differentiable Consistency Constraints for Improved Deep Speech Enhancement
复制标题
用于改进深度语音增强的可微一致性约束
DOI:
10.1109/icassp.2019.8682783
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
R. Saurous
中科院分区:
文献类型:
--
作者:
Scott Wisdom;J. Hershey;K. Wilson;J. Thorpe;Michael Chinen;Brian Patton;R. Saurous
In recent years, deep networks have led to dramatic improvements in speech enhancement by framing it as a data-driven pattern recognition problem. In many modern enhancement systems, large amounts of data are used to train a deep network to estimate masks for complex-valued short-time Fourier transforms (STFTs) to suppress noise and preserve speech. However, current masking approaches often neglect two important constraints: STFT consistency and mixture consistency. Without STFT consistency, the system’s output is not necessarily the STFT of a time-domain signal, and without mixture consistency, the sum of the estimated sources does not necessarily equal the input mixture. Furthermore, the only previous approaches that apply mixture consistency use real-valued masks; mixture consistency has been ignored for complex-valued masks.