A near-end listening enhancement system by RNN-based noise cancellation and speech modification

A near-end listening enhancement system by RNN-based noise cancellation and speech modification
复制标题

基于 RNN 的噪声消除和语音修改的近端听力增强系统

DOI:
10.1007/s11042-018-6947-8
复制
发表时间:
2019
影响因子:
3.6
通讯作者:
Zhang Rui
Zhang Rui
中科院分区:
计算机科学4区
文献类型:
--
作者:
Li Gang;Hu Ruimin;Wang Xiaochen;Zhang Rui

文献摘要

相似文献

当人们在嘈杂的环境中听电话时,近端听力增强(NEIL)是一种提高语音清晰度对抗环境噪声的技术。移动通信中的复杂环境激发了许多学者对NELE的研究。虽然他们提出了许多NELEL系统,但他们只关注语音的修改,以提高可理解性。很少有学者尝试通过噪声消除来进一步提高可理解性。因为传统的噪声消除是基于自适应滤波的。如果在最常见的手机模式下使用自适应滤波,由于反馈麦克风暴露在复杂的环境中,导致反馈不充分,噪声消除效果会很差。随着深度神经网络(DNN)的蓬勃发展,DNN能够在没有反馈麦克风的情况下对噪声信号进行预测以消除噪声,特别是对于递归神经网络(RNN)。在这项研究中,我们提出了一种基于RNN的噪声消除和语音修改的NEELE系统(RNC-SM),它在语音修改后引入了噪声消除功能。与现有的NELEL系统相比,RNC-SM系统有效地提高了客观语音清晰度指数(SII)得分和主观收听质量。
When people listen to the phone in noisy environments, near-end listening enhancement (NELE) is a technology to enhance speech intelligibility against environmental noise. The complex environments in mobile communications have inspired many scholars to engage in NELE researches. Although they have proposed a lot of NELE systems, they only focus on the speech modification to enhance the intelligibility. Few scholars have attempted to further enhance the intelligibility by noise cancellation. Because traditional noise cancellation is based on adaptive filtering. If the adaptive filtering is used in the most common handset mode, the noise cancellation result will be poor because of inadequate feedback caused by the feedback microphone exposed to complex environments. With the booming of the deep neural network (DNN), DNN is able to predict noise signals for noise cancellation without the feedback microphone, especially for recurrent neural network (RNN). In this study, we propose a NELE System by RNN-based noise cancellation and speech modification (RNC-SM), which introduce a noise cancellation function after speech modification. Compared with existing NELE systems, RNC-SM system effectively improves the objective speech intelligibility index (SII) scores and the subjective listening quality.