Source counting in speech mixtures using a variational EM approach for complex WATSON mixture models

Source counting in speech mixtures using a variational EM approach for complex WATSON mixture models
复制标题

DOI:
10.1109/icassp.2014.6854924
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Lukas Drude;Aleksej Chinaev;Dang Hai Tran Vu;Reinhold Häb-Umbach
Lukas Drude;Aleksej Chinaev;Dang Hai Tran Vu;Reinhold Häb-Umbach
中科院分区:
其他
文献类型:
--
作者:
Lukas Drude;Aleksej Chinaev;Dang Hai Tran Vu;Reinhold Häb-Umbach

文献摘要

被引文献

相似文献

在这方面的贡献,我们推导出一个变分EM(VEM)算法在复杂的沃森混合模型,最近已被提出作为一个模型的归一化麦克风阵列信号的分布在短时傅立叶变换域中的模型选择。VEM算法通过迭代估计沃森分布的模式向量并抑制来自相应方向的信号来计算混合语音中的活动源的数量。一个关键的理论贡献是推导的MMSE估计的二次型涉及的模式向量的沃森分布。实验结果表明,在中等低信噪比下,信源计数方法是有效的。它进一步表明,VEM算法是更强大的使用阈值。
In this contribution we derive a variational EM (VEM) algorithm for model selection in complex Watson mixture models, which have been recently proposed as a model of the distribution of normalized microphone array signals in the short-time Fourier transform domain. The VEM algorithm is applied to count the number of active sources in a speech mixture by iteratively estimating the mode vectors of the Watson distributions and suppressing the signals from the corresponding directions. A key theoretical contribution is the derivation of the MMSE estimate of a quadratic form involving the mode vector of the Watson distribution. The experimental results demonstrate the effectiveness of the source counting approach at moderately low SNR. It is further shown that the VEM algorithm is more robust with respect to used threshold values.