Phasebook and Friends: Leveraging Discrete Representations for Source Separation

Phasebook and Friends: Leveraging Discrete Representations for Source Separation
复制标题

DOI:
10.1109/jstsp.2019.2904183
复制
发表时间:
2018-10
影响因子:
7.5
通讯作者:
J. Le Roux;G. Wichern;Shinji Watanabe;Andy M. Sarroff;J. Hershey
J. Le Roux;G. Wichern;Shinji Watanabe;Andy M. Sarroff;J. Hershey
中科院分区:
工程技术1区
文献类型:
--
作者:
J. Le Roux;G. Wichern;Shinji Watanabe;Andy M. Sarroff;J. Hershey

文献摘要

被引文献

相似文献

基于深度学习的语音增强和源分离系统最近达到了前所未有的质量水平,以至于性能达到了一个新的上限。大多数系统依赖于估计目标源的幅度,通过估计一个实值掩模来应用于混合信号的时频表示。这种方法的一个限制因素是缺乏相位估计:在重建估计的时域信号时最常使用混合的相位。在这里,我们提出了“magbook”、“phasbook”和“combook”这三种基于离散表示的新型层,可用于估计复杂的时频掩模。Magbook层扩展了经典的s型单元和最近引入的凸softmax激活,用于基于掩模的震级估计。相簿层使用类似的结构来给出相位掩模的估计,而不会受到相位包裹问题的困扰。Combook层是magbook-phasebook组合的另一种选择,可以直接估计复杂的掩模。我们提出了涉及这些表示的各种训练和推理方案,并特别解释了如何将它们包含在端到端学习框架中。我们还提出了一项oracle研究,以评估使用离散相位表示的各种类型掩模的性能上限。我们在wsj0-2mix数据集上评估了所提出的方法,该数据集是一个经过充分研究的用于单通道扬声器独立扬声器分离的语料库,其性能与最先进的基于掩模的方法相匹配,而无需额外的相位重建步骤。
Speech enhancement and source separation systems based on deep learning have recently reached unprecedented levels of quality, to the point that performance is reaching a new ceiling. Most systems rely on estimating the magnitude of a target source by estimating a real-valued mask to be applied to a time-frequency representation of the mixture signal. A limiting factor in such approaches is a lack of phase estimation: the phase of the mixture is most often used when reconstructing the estimated time-domain signal. Here, we propose “magbook,” “phasebook,” and “combook,” three new types of layers based on discrete representations that can be used to estimate complex time-frequency masks. Magbook layers extend classical sigmoidal units and a recently introduced convex softmax activation for mask-based magnitude estimation. Phasebook layers use a similar structure to give an estimate of the phase mask without suffering from phase wrapping issues. Combook layers are an alternative to the magbook–phasebook combination that directly estimate complex masks. We present various training and inference schemes involving these representations, and explain in particular how to include them in an end-to-end learning framework. We also present an oracle study to assess upper bounds on performance for various types of masks using discrete phase representations. We evaluate the proposed methods on the wsj0-2mix dataset, a well-studied corpus for single-channel speaker-independent speaker separation, matching the performance of state-of-the-art mask-based approaches without requiring additional phase reconstruction steps.