Neural Full-Rank Spatial Covariance Analysis for Blind Source Separation

Neural Full-Rank Spatial Covariance Analysis for Blind Source Separation
复制标题

盲源分离的神经全秩空间协方差分析

DOI:
10.1109/lsp.2021.3101699
复制
发表时间:
2021
影响因子:
3.9
通讯作者:
Kazuyoshi Yoshii
Kazuyoshi Yoshii
中科院分区:
工程技术2区
文献类型:
--
作者:
Yoshiaki Bando;Kouhei Sekiguchi;Yoshiki Masuyama;Aditya Arie Nugraha;Mathieu Fontaine;Kazuyoshi Yoshii

文献摘要

被引文献

相似文献

本文提出了一种基于混合信号非线性生成模型的神经网络盲源分离方法。盲源分离的一个经典统计方法是拟合线性生成模型,该模型由空间模型和源模型组成,分别表示源的通道间协方差和功率谱密度。虽然变分自动编码器(VAE)已经被成功地用作具有潜在特征的非线性信源模型,但它应该从足够数量的孤立信号中进行预训练。相反,我们的方法使得基于VAE的源模型能够仅从混合信号中训练。具体地说,我们引入了一种神经混合-特征推理模型,该模型直接从观察到的混合中推断出潜在的特征,并将其与由满阶空间模型和基于VAE的信源模型组成的神经特征-混合生成模型相结合。所有模型都被联合优化,使得训练混合的可能性在AVI框架下最大化。一旦推理模型被优化,它就可以用于估计包含在不可见混合信号中的源的潜在特征。实验结果表明,该方法的性能优于现有的基于线性生成模型的盲源分离方法,与基于VAE源模型的监督学习方法相当。
This paper describes aneural blind source separation (BSS) method based on amortized variational inference (AVI) of a non-linear generative model of mixture signals. A classical statistical approach to BSS is to fit a linear generative model that consists of spatial and source models representing the inter-channel covariances and power spectral densities of sources, respectively. Although the variational autoencoder (VAE) has successfully been used as a non-linear source model with latent features, it should be pretrained from a sufficient amount of isolated signals. Our method, in contrast, enables the VAE-based source model to be trained only from mixture signals. Specifically, we introduce a neural mixture-to-feature inference model that directly infers the latent features from the observed mixture and integrate it with a neural feature-to-mixture generative model consisting of a full-rank spatial model and a VAE-based source model. All the models are optimized jointly such that the likelihood for the training mixtures is maximized in the framework of AVI. Once the inference model is optimized, it can be used for estimating the latent features of sources included in unseen mixture signals. The experimental results show that the proposed method outperformed the state-of-the-art BSS methods based on linear generative models and was comparable to a method based on supervised learning of the VAE-based sourcemodel.