Use of bimodal coherence to resolve the permutation problem in convolutive BSS

Use of bimodal coherence to resolve the permutation problem in convolutive BSS
复制标题

DOI:
10.1016/j.sigpro.2011.11.007
复制
发表时间:
2012-08
期刊:
Signal Process.
影响因子:
--
通讯作者:
Qingju Liu;Wenwu Wang;P. Jackson
Qingju Liu;Wenwu Wang;P. Jackson
中科院分区:
其他
文献类型:
--
作者:
Qingju Liu;Wenwu Wang;P. Jackson

文献摘要

被引文献

相似文献

最近的研究表明,视觉语音中包含的面部信息有助于提高纯音频盲源分离(BSS)算法的性能。通过使用例如高斯混合模型(GMM)对音频和可视语音之间的相干性进行统计表征来利用这种信息。在本文中,我们提出了三点贡献。利用同步特征,我们提出了一种自适应期望最大化(AEM)算法来对离线训练过程中的视听一致性进行建模。为了提高该相干模型的精度,我们使用了一种帧选择方案来丢弃非平稳特征。然后,利用相干最大化技术,提出了一种新的排序方法来解决频域中的排列问题。我们在一个由元音和辅音的不同组合组成的多模式语音数据库上测试了我们的算法。实验结果表明,该算法的性能优于传统的纯音频盲源分离算法,证实了利用视觉语音辅助音频分离的优点。
Recent studies show that facial information contained in visual speech can be helpful for the performance enhancement of audio-only blind source separation (BSS) algorithms. Such information is exploited through the statistical characterization of the coherence between the audio and visual speech using, e.g., a Gaussian mixture model (GMM). In this paper, we present three contributions. With the synchronized features, we propose an adapted expectation maximization (AEM) algorithm to model the audio–visual coherence in the off-line training process. To improve the accuracy of this coherence model, we use a frame selection scheme to discard nonstationary features. Then with the coherence maximization technique, we develop a new sorting method to solve the permutation problem in the frequency domain. We test our algorithm on a multimodal speech database composed of different combinations of vowels and consonants. The experimental results show that our proposed algorithm outperforms traditional audio-only BSS, which confirms the benefit of using visual speech to assist in separation of the audio.