Fast Multichannel Nonnegative Matrix Factorization With Directivity-Aware Jointly-Diagonalizable Spatial Covariance Matrices for Blind Source Separation

Fast Multichannel Nonnegative Matrix Factorization With Directivity-Aware Jointly-Diagonalizable Spatial Covariance Matrices for Blind Source Separation
复制标题

DOI:
10.1109/taslp.2020.3019181
复制
发表时间:
2020-08
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Kouhei Sekiguchi;Yoshiaki Bando;Aditya Arie Nugraha;Kazuyoshi Yoshii;Tatsuya Kawahara
Kouhei Sekiguchi;Yoshiaki Bando;Aditya Arie Nugraha;Kazuyoshi Yoshii;Tatsuya Kawahara
中科院分区:
其他
文献类型:
--
作者:
Kouhei Sekiguchi;Yoshiaki Bando;Aditya Arie Nugraha;Kazuyoshi Yoshii;Tatsuya Kawahara

文献摘要

相似文献

本文描述了一种计算效率高的盲源分离(BSS)方法的基础上的独立性,低秩,和方向性的来源。BSS的典型方法是概率模型的无监督学习,该概率模型由表示源图像的时频结构的源模型和表示它们的通道间协方差结构的空间模型组成。基于非负矩阵分解(NMF)的低秩源模型被认为是频间源对齐的有效方法,多通道NMF(MNMF)假设源图像服从多元复高斯分布,空间协方差矩阵(SCM)不受约束。限制SCM的自由度是降低MNMF计算量和初始化敏感性的有效途径。虽然MNMF的一个变种称为独立低秩矩阵分析(ILRMA)严格限制SCM秩1矩阵的理想化条件下,只有方向性和低回声源存在,我们限制SCM联合对角化,但满秩矩阵在频率方面的方式,导致FastMNMF1。为了帮助频率间源对齐,我们提出了FastMNMF 2,它在所有频率区间上共享每个源的方向特征。为了明确考虑每个源的方向性或扩散性,我们还提出了秩约束FastMNMF,使我们能够单独指定SCM的秩。实验结果表明,FastMNMF在语音分离方面优于MNMF和ILRMA,并且在语音增强中,秩约束是有效的。
This article describes a computationally-efficient blind source separation (BSS) method based on the independence, low-rankness, and directivity of the sources. A typical approach to BSS is unsupervised learning of a probabilistic model that consists of a source model representing the time-frequency structure of source images and a spatial model representing their inter-channel covariance structure. Building upon the low-rank source model based on nonnegative matrix factorization (NMF), which has been considered to be effective for inter-frequency source alignment, multichannel NMF (MNMF) assumes source images to follow multivariate complex Gaussian distributions with unconstrained full-rank spatial covariance matrices (SCMs). An effective way of reducing the computational cost and initialization sensitivity of MNMF is to restrict the degree of freedom of SCMs. While a variant of MNMF called independent low-rank matrix analysis (ILRMA) severely restricts SCMs to rank-1 matrices under an idealized condition that only directional and less-echoic sources exist, we restrict SCMs to jointly-diagonalizable yet full-rank matrices in a frequency-wise manner, resulting in FastMNMF1. To help inter-frequency source alignment, we then propose FastMNMF2 that shares the directional feature of each source over all frequency bins. To explicitly consider the directivity or diffuseness of each source, we also propose rank-constrained FastMNMF that enables us to individually specify the ranks of SCMs. Our experiments showed the superiority of FastMNMF over MNMF and ILRMA in speech separation and the effectiveness of the rank constraint in speech enhancement.