Long short-term memory for speaker generalization in supervised speech separation

Long short-term memory for speaker generalization in supervised speech separation
复制标题

DOI:
10.1121/1.4986931
复制
发表时间:
2017-06-01
影响因子:
2.4
通讯作者:
Wang, DeLiang
Wang, DeLiang
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Chen, Jitong;Wang, DeLiang

文献摘要

被引文献

相似文献

语音分离可以用公式表示为学习从噪声语音中提取的声学特征估计时频掩模。对于有监督的语音分离,泛化到看不见的噪声和看不见的说话人是一个关键问题。尽管深度神经网络(DNN)在与噪声无关的语音分离方面取得了成功,但DNN在对大量说话者建模方面受到限制。为了提高说话人泛化能力,提出了一种基于长短期记忆(LSTM)的分离模型,该模型自然地考虑了语音的时间动态。系统的评估表明,该模型大大优于基于DNN的模型看不见的扬声器和看不见的噪声在客观语音清晰度。分析LSTM内部表示可以发现,LSTM捕获了长期的语音上下文。还发现LSTM模型对于低延迟语音分离更有利,并且在没有未来帧的情况下,它比具有未来帧的DNN模型表现得更好。该模型是一种有效的方法,说话人和噪声无关的语音分离。(C)2017年美国声学学会。
Speech separation can be formulated as learning to estimate a time-frequency mask from acoustic features extracted from noisy speech. For supervised speech separation, generalization to unseen noises and unseen speakers is a critical issue. Although deep neural networks (DNNs) have been successful in noise-independent speech separation, DNNs are limited in modeling a large number of speakers. To improve speaker generalization, a separation model based on long short-term memory (LSTM) is proposed, which naturally accounts for temporal dynamics of speech. Systematic evaluation shows that the proposed model substantially outperforms a DNN-based model on unseen speakers and unseen noises in terms of objective speech intelligibility. Analyzing LSTM internal representations reveals that LSTM captures long-term speech contexts. It is also found that the LSTM model is more advantageous for low-latency speech separation and it, without future frames, performs better than the DNN model with future frames. The proposed model represents an effective approach for speaker-and noise-independent speech separation. (C) 2017 Acoustical Society of America.