Joint Mixing Vector and Binaural Model Based Stereo Source Separation

Joint Mixing Vector and Binaural Model Based Stereo Source Separation
复制标题

DOI:
10.1109/taslp.2014.2320637
复制
发表时间:
2014-09
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Atiyeh Alinaghi;P. Jackson;Qingju Liu;Wenwu Wang
Atiyeh Alinaghi;P. Jackson;Qingju Liu;Wenwu Wang
中科院分区:
其他
文献类型:
--
作者:
Atiyeh Alinaghi;P. Jackson;Qingju Liu;Wenwu Wang

文献摘要

被引文献

相似文献

在本文中,混合向量(MV)的统计混合模型进行了比较双耳线索所表示的两耳电平和相位差(ILD和IPD)。结果表明,MV分布是相当不同的,而双耳模型重叠时,源是彼此接近。另一方面,双耳线索比MV模型对高混响更鲁棒。根据这种互补的行为,我们介绍了一种新的立体声语音分离的鲁棒算法,该算法同时考虑了加性和卷积噪声信号,并行地对MV和双耳线索进行建模,并估计概率时频掩模。每个线索的最终决定的贡献也调整加权对数似然的线索经验。此外,频域盲源分离(BSS)的置换问题,解决了初始化MV的双耳线索的基础上。实验系统地进行了确定和欠定语音混合在五个房间的各种声学特性,包括消声,高混响,空间扩散噪声条件。在信号失真比(SDR)方面的结果证实了整合MV和双耳线索的好处,相比,两个国家的最先进的基线算法,只使用MV或双耳线索。
In this paper the mixing vector (MV) in the statistical mixing model is compared to the binaural cues represented by interaural level and phase differences (ILD and IPD). It is shown that the MV distributions are quite distinct while binaural models overlap when the sources are close to each other. On the other hand, the binaural cues are more robust to high reverberation than MV models. According to this complementary behavior we introduce a new robust algorithm for stereo speech separation which considers both additive and convolutive noise signals to model the MV and binaural cues in parallel and estimate probabilistic time-frequency masks. The contribution of each cue to the final decision is also adjusted by weighting the log-likelihoods of the cues empirically. Furthermore, the permutation problem of the frequency domain blind source separation (BSS) is addressed by initializing the MVs based on binaural cues. Experiments are performed systematically on determined and underdetermined speech mixtures in five rooms with various acoustic properties including anechoic, highly reverberant, and spatially-diffuse noise conditions. The results in terms of signal-to-distortion-ratio (SDR) confirm the benefits of integrating the MV and binaural cues, as compared with two state-of-the-art baseline algorithms which only use MV or the binaural cues.