One-Shot Conditional Audio Filtering of Arbitrary Sounds

One-Shot Conditional Audio Filtering of Arbitrary Sounds
复制标题

任意声音的一次性条件音频过滤

DOI:
--
复制
发表时间:
2020
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
M. Tagliasacchi
M. Tagliasacchi
中科院分区:
--
文献类型:
--
作者:
Beat Gfeller;Dominik Roblek;M. Tagliasacchi

文献摘要

参考文献

被引文献

相似文献

我们考虑仅基于目标源的简短样本(来自同一录音)从单通道混合中分离特定声源的问题。使用 SoundFilter(一种波到波神经网络架构),我们可以在不使用任何声音类别标签的情况下训练模型。使用与源分离网络联合学习的调节编码器模型,可以“配置”训练后的模型来过滤任意声源,甚至是在训练期间未见过的声源。在 FSD50k 数据集上进行评估,我们的模型对于两种声音的混合获得了 9.6 dB 的 SI-SDR 改进。在 Librispeech 上进行训练时,我们的模型在从两个说话者的混合声音中分离出一种声音时,SI-SDR 提高了 14.0 dB。此外,我们表明,条件编码器学习到的表示在嵌入空间中将声学上相似的声音聚集在一起,即使它是在不使用任何标签的情况下进行训练的。
We consider the problem of separating a particular sound source from a single-channel mixture, based on only a short sample of the target source (from the same recording). Using SoundFilter, a wave-to-wave neural network architecture, we can train a model without using any sound class labels. Using a conditioning encoder model which is learned jointly with the source separation network, the trained model can be "configured" to filter arbitrary sound sources, even ones that it has not seen during training. Evaluated on the FSD50k dataset, our model obtains an SI-SDR improvement of 9.6 dB for mixtures of two sounds. When trained on Librispeech, our model achieves an SI-SDR improvement of 14.0 dB when separating one voice from a mixture of two speakers. Moreover, we show that the representation learned by the conditioning encoder clusters acoustically similar sounds together in the embedding space, even though it is trained without using any labels.
深度(呃)学习。
DOI: 10.1523/jneurosci.0153-18.2018
发表时间: 2018
期刊: The Journal of neuroscience : the official journal of the Society for Neuroscience
影响因子: --
作者:
Srinivasan,Shyam;Greenspan,RalphJ;Stevens,CharlesF;Grover,Dhruv
通讯作者: Grover,Dhruv