Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices

Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices
复制标题

基于联合对角空间协方差矩阵的快速多通道源分离

DOI:
10.23919/eusipco.2019.8902557
复制
发表时间:
2019
期刊:
2019 27th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
Kazuyoshi Yoshii
Kazuyoshi Yoshii
中科院分区:
--
文献类型:
--
作者:
Kouhei Sekiguchi;A. Nugraha;Yoshiaki Bando;Kazuyoshi Yoshii

文献摘要

参考文献

被引文献

相似文献

本文描述了一种通用的方法,该方法加速了基于满阶空间建模的多通道源分离方法。一种流行的多声道声源分离方法是将空间模型与声源模型相结合,在时频域中估计每个声源的空间协方差矩阵(SCM)和功率谱密度(PSD)。这种方法最成功的例子之一是基于全秩空间模型和低秩源模型的多通道非负矩阵分解(MNMF)。然而,MNMF的计算代价很高,而且由于很难估计无约束满阶SCM,因此往往效果不佳。我们不像在独立低阶矩阵分析(ILRMA)中那样将SCM限制为具有严重丧失空间建模能力的1阶矩阵,而是将每个频率段的SCM限制为可联合对角化但仍然是满秩矩阵。对于这种快速版本的MNMF,我们提出了一种在形式上类似于ILRMA的计算效率和收敛保证的算法。类似地,我们提出了一种基于深度语音模型和低阶噪声模型的最新语音增强方法的快速版本。实验结果表明,MNMF的快速版本和深度语音增强方法的速度分别是原始版本的几倍,甚至更好。
This paper describes a versatile method that accelerates multichannel source separation methods based on full-rank spatial modeling. A popular approach to multichannel source separation is to integrate a spatial model with a source model for estimating the spatial covariance matrices (SCMs) and power spectral densities (PSDs) of each sound source in the time-frequency domain. One of the most successful examples of this approach is multichannel nonnegative matrix factorization (MNMF) based on a full-rank spatial model and a low-rank source model. MNMF, however, is computationally expensive and often works poorly due to the difficulty of estimating the unconstrained full-rank SCMs. Instead of restricting the SCMs to rank -1 matrices with the severe loss of the spatial modeling ability as in independent low-rank matrix analysis (ILRMA), we restrict the SCMs of each frequency bin to jointly-diagonalizable but still full-rank matrices. For such a fast version of MNMF, we propose a computationally-efficient and convergence-guaranteed algorithm that is similar in form to that of ILRMA. Similarly, we propose a fast version of a state of-the-art speech enhancement method based on a deep speech model and a low-rank noise model. Experimental results showed that the fast versions of MNMF and the deep speech enhancement method were several times faster and performed even better than the original versions of those methods, respectively.
DOI: 10.1016/j.csl.2016.10.005
发表时间: 2017-11-01
影响因子: 4.3
作者:
Barker, Jon;Marxer, Ricard;Watanabe, Shinji
通讯作者: Watanabe, Shinji