Neural Cascade Architecture for Multi-Channel Acoustic Echo Suppression

Neural Cascade Architecture for Multi-Channel Acoustic Echo Suppression
复制标题

DOI:
10.1109/taslp.2022.3192104
复制
发表时间:
2022
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Zhang;Deliang Wang
H. Zhang;Deliang Wang
中科院分区:
其他
文献类型:
--
作者:
H. Zhang;Deliang Wang

文献摘要

相似文献

传统的回声抵消(AEC)通过使用自适应算法识别声脉冲响应来工作。为了解决单通道和多通道AEC(MCAEC)问题,提出了一种用于声回波和噪声联合抑制的神经级联结构。建议的级联体系结构由两个模块组成。在第一个模块中使用卷积递归网络(CRN)进行复谱映射。其输出作为附加输入馈送到第二模块,其中利用长短期记忆网络(LSTM)进行幅度掩码估计。整个体系结构以端到端的方式进行训练,两个模块使用单个损耗函数进行联合优化。使用分别从第一和第二模块获得的增强的相位和幅度来生成最终输出。级联结构使所提出的方法能够获得稳健的幅度估计和相位增强。在不同的AEC设置下对所提出的方法进行了研究。我们发现,基于深度学习的方法避免了传统MCAEC中的非唯一性问题。对于具有多个麦克风的MCAEC设置,将深度MCAEC与监督波束形成相结合进一步提高了系统性能。评估结果表明,该方法在保持语音质量的同时,有效地抑制了声学回波和噪声,并且在不同设置下的性能一致优于相关方法。
Traditional acoustic echo cancellation (AEC) works by identifying an acoustic impulse response using adaptive algorithms. This paper proposes a neural cascade architecture for joint acoustic echo and noise suppression to address both single-channel and multi-channel AEC (MCAEC) problems. The proposed cascade architecture consists of two modules. A convolutional recurrent network (CRN) is employed in the first module for complex spectral mapping. Its output is fed as an additional input to the second module, where a long short-term memory network (LSTM) is utilized for magnitude mask estimation. The entire architecture is trained in an end-to-end manner with the two modules optimized jointly using a single loss function. The final output is generated using the enhanced phase and magnitude obtained from the first and the second module, respectively. The cascade architecture enables the proposed method to obtain robust magnitude estimation as well as phase enhancement. The proposed method is investigated under different AEC setups. We find that the deep learning based approach avoids the no-uniqueness problem in traditional MCAEC. For MCAEC setups with multiple microphones, combining deep MCAEC with supervised beamforming further improves the system performance. Evaluation results show that the proposed approach effectively suppresses acoustic echo and noise while preserving speech quality, and consistently outperforms related methods under different setups.