Unsupervised Sound Separation Using Mixture Invariant Training

Unsupervised Sound Separation Using Mixture Invariant Training
复制标题

使用混合不变训练的无监督声音分离

DOI:
--
复制
发表时间:
2020
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
J. Hershey
J. Hershey
中科院分区:
--
文献类型:
--
作者:
Scott Wisdom;Efthymios Tzinis;Hakan Erdogan;Ron J. Weiss;K. Wilson;J. Hershey

文献摘要

被引文献

相似文献

近年来,使用深度神经网络的监督训练在单通道声音分离问题上取得了快速进展。在这种监督方法中,模型被训练以从通过将孤立的地面实况源相加而创建的合成混合物中预测分量源。依赖这种合成训练数据是有问题的,因为良好的性能取决于训练数据与真实世界音频之间的匹配程度,特别是在声学条件和源分布方面。声学特性可能难以准确模拟,并且声音类型的分布可能难以复制。在本文中,我们提出了一种完全无监督的方法,混合不变训练(MixIT),它只需要单通道声学混合。在MixIT中,通过将现有混合物混合在一起来构建训练示例,并且模型将它们分离成可变数量的潜在源,以便可以重新混合分离的源以近似原始混合物。我们表明,MixIT可以实现有竞争力的性能相比,监督的语音分离方法。在半监督学习设置中使用MixIT可以实现无监督域自适应和从大量真实的世界数据中学习,而无需地面实况源波形。特别是,我们显着提高混响语音分离的性能,通过将混响混合,训练语音增强系统从嘈杂的混合物,并提高通用的声音分离,结合大量的野外数据。
In recent years, rapid progress has been made on the problem of single-channel sound separation using supervised training of deep neural networks. In such supervised approaches, a model is trained to predict the component sources from synthetic mixtures created by adding up isolated ground-truth sources. Reliance on this synthetic training data is problematic because good performance depends upon the degree of match between the training data and real-world audio, especially in terms of the acoustic conditions and distribution of sources. The acoustic properties can be challenging to accurately simulate, and the distribution of sound types may be hard to replicate. In this paper, we propose a completely unsupervised method, mixture invariant training (MixIT), that requires only single-channel acoustic mixtures. In MixIT, training examples are constructed by mixing together existing mixtures, and the model separates them into a variable number of latent sources, such that the separated sources can be remixed to approximate the original mixtures. We show that MixIT can achieve competitive performance compared to supervised methods on speech separation. Using MixIT in a semi-supervised learning setting enables unsupervised domain adaptation and learning from large amounts of real world data without ground-truth source waveforms. In particular, we significantly improve reverberant speech separation performance by incorporating reverberant mixtures, train a speech enhancement system from noisy mixtures, and improve universal sound separation by incorporating a large amount of in-the-wild data.