DNN-free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online Fastmnmf

DNN-free Low-Latency Adaptive Speech Enhancement Based on Frame-Online Beamforming Powered by Block-Online Fastmnmf
复制标题

DOI:
10.1109/iwaenc53105.2022.9914729
复制
发表时间:
2022-07
期刊:
2022 International Workshop on Acoustic Signal Enhancement (IWAENC)
影响因子:
--
通讯作者:
Aditya Arie Nugraha;Kouhei Sekiguchi;Mathieu Fontaine;Yoshiaki Bando;Kazuyoshi Yoshii
Aditya Arie Nugraha;Kouhei Sekiguchi;Mathieu Fontaine;Yoshiaki Bando;Kazuyoshi Yoshii
中科院分区:
其他
文献类型:
--
作者:
Aditya Arie Nugraha;Kouhei Sekiguchi;Mathieu Fontaine;Yoshiaki Bando;Kazuyoshi Yoshii

文献摘要

相似文献

本文描述了一种实用的双处理语音增强系统,该系统采用环境敏感的帧在线波束形成(前端),并辅以无环境的块在线源分离(后端)。为了使用最小方差无失真响应(MVDR)波束成形,可以训练深度神经网络(DNN),该深度神经网络估计用于计算源(语音和噪声)的协方差矩阵的时频掩模。提出了一种基于反向传播的DNN运行时自适应算法,用于处理训练-测试条件不匹配的情况。相反,人们可以尝试直接估计源协方差矩阵与国家的最先进的盲源分离方法称为快速多通道非负矩阵分解(FastMNMF)。然而,在实践中,DNN和FastMNMF都不能以帧在线方式更新,这是由于其计算昂贵的迭代性质。我们的无DNN系统利用块在线FastMNMF给出的最新源频谱图的后验来推导当前源协方差矩阵,用于帧在线波束成形。评估表明,我们的帧在线系统可以快速响应由干扰扬声器移动引起的场景变化,并且在字错误率方面优于现有的基于DNN的波束成形的块在线系统5.0个点。
This paper describes a practical dual-process speech enhancement system that adapts environment-sensitive frame-online beamforming (front-end) with help from environment-free block-online source separation (back-end). To use minimum variance distortionless response (MVDR) beamforming, one may train a deep neural network (DNN) that estimates time-frequency masks used for computing the covariance matrices of sources (speech and noise). Backpropagation-based run-time adaptation of the DNN was proposed for dealing with the mismatched training-test conditions. Instead, one may try to directly estimate the source covariance matrices with a state-of-the-art blind source separation method called fast multichannel non-negative matrix factorization (FastMNMF). In practice, however, neither the DNN nor the FastMNMF can be updated in a frame-online manner due to its computationally-expensive iterative nature. Our DNN-free system leverages the posteri-ors of the latest source spectrograms given by block-online FastMNMF to derive the current source covariance matrices for frame-online beamforming. The evaluation shows that our frame-online system can quickly respond to scene changes caused by interfering speaker movements and outperformed an existing block-online system with DNN-based beamforming by 5.0 points in terms of the word error rate.