Listening to Sounds of Silence for Speech Denoising

Listening to Sounds of Silence for Speech Denoising
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Ruilin Xu;Rundi Wu;Y. Ishiwaka;Carl Vondrick;Changxi Zheng
Ruilin Xu;Rundi Wu;Y. Ishiwaka;Carl Vondrick;Changxi Zheng
中科院分区:
其他
文献类型:
--
作者:
Ruilin Xu;Rundi Wu;Y. Ishiwaka;Carl Vondrick;Changxi Zheng

文献摘要

被引文献

相似文献

我们介绍了语音去噪的深度学习模型,语音去噪是音频分析中许多应用中出现的长期挑战。我们的方法是基于对人类语言的一个关键观察:每个句子或单词之间通常会有一个短暂的停顿。在录制的语音信号中,这些停顿引入了一系列只有噪声存在的时间段。我们利用这些偶然的沉默间隔来学习一个仅给定单声道音频的自动语音去噪模型。随着时间的推移,检测到的沉默间隔不仅暴露了纯噪声,而且暴露了其时变特征,使模型能够学习噪声动态并从语音信号中抑制噪声。在多个数据集上的实验证实了无声间隔检测对语音去噪的关键作用,我们的方法优于几种最先进的去噪方法,包括那些只接受音频输入(如我们的)和那些基于视听输入(因此需要更多信息)去噪的方法。我们还表明,我们的方法具有出色的泛化特性,例如去噪训练中未见过的口语。
We introduce a deep learning model for speech denoising, a long-standing challenge in audio analysis arising in numerous applications. Our approach is based on a key observation about human speech: there is often a short pause between each sentence or word. In a recorded speech signal, those pauses introduce a series of time periods during which only noise is present. We leverage these incidental silent intervals to learn a model for automatic speech denoising given only mono-channel audio. Detected silent intervals over time expose not just pure noise but its time-varying features, allowing the model to learn noise dynamics and suppress it from the speech signal. Experiments on multiple datasets confirm the pivotal role of silent interval detection for speech denoising, and our method outperforms several state-of-the-art denoising methods, including those that accept only audio input (like ours) and those that denoise based on audiovisual input (and hence require more information). We also show that our method enjoys excellent generalization properties, such as denoising spoken languages not seen during training.