Underdetermined Convolutive Blind Source Separation via Frequency Bin-Wise Clustering and Permutation Alignment

Underdetermined Convolutive Blind Source Separation via Frequency Bin-Wise Clustering and Permutation Alignment
复制标题

DOI:
10.1109/tasl.2010.2051355
复制
发表时间:
2011-03
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Sawada;S. Araki;S. Makino
H. Sawada;S. Araki;S. Makino
中科院分区:
其他
文献类型:
--
作者:
H. Sawada;S. Araki;S. Makino

文献摘要

被引文献

相似文献

本文提出了一种语音/音频卷积混合信号的盲源分离方法。该方法甚至可以应用到一个欠定的情况下,有更少的麦克风比源。分离操作在频域中执行,并且由两个阶段组成。在第一阶段,频域混合样本聚类到每个源的期望最大化(EM)算法。由于聚类是以频率分箱方式执行的,因此分箱聚类样本的排列模糊度应该对齐。这在第二阶段通过使用每个样本属于指定类的可能性的概率来解决。这种两级结构使得即使在混响条件下也可以实现良好的分离。在混响条件下用三个麦克风分离四个语音信号的实验结果表明,新方法优于现有的方法。我们还报告分离结果的基准数据集和现场录音的语音混合物。
This paper presents a blind source separation method for convolutive mixtures of speech/audio sources. The method can even be applied to an underdetermined case where there are fewer microphones than sources. The separation operation is performed in the frequency domain and consists of two stages. In the first stage, frequency-domain mixture samples are clustered into each source by an expectation-maximization (EM) algorithm. Since the clustering is performed in a frequency bin-wise manner, the permutation ambiguities of the bin-wise clustered samples should be aligned. This is solved in the second stage by using the probability on how likely each sample belongs to the assigned class. This two-stage structure makes it possible to attain a good separation even under reverberant conditions. Experimental results for separating four speech signals with three microphones under reverberant conditions show the superiority of the new method over existing methods. We also report separation results for a benchmark data set and live recordings of speech mixtures.