Fusing audio and visual features of speech

Fusing audio and visual features of speech
复制标题

融合语音的音频和视觉特征

DOI:
10.1109/icip.2000.899333
复制
发表时间:
2000
期刊:
Proceedings 2000 International Conference on Image Processing (Cat. No.00CH37101)
影响因子:
--
通讯作者:
Thomas S. Huang
Thomas S. Huang
中科院分区:
--
文献类型:
--
作者:
Hao Pan;Zhi;Thomas S. Huang

文献摘要

被引文献

相似文献

本文采用一种新的融合隐马尔可夫模型,将语音的视听特征融合在一起。我们假设这两组特征可以具有不同的数据速率和持续时间。首先利用隐马尔可夫模型(HMM)分别对它们进行建模,然后利用在最大熵意义下最优的广义贝叶斯融合方法将它们融合在一起。特别介绍了一种高效的学习算法。该学习算法不是最大化融合-隐马尔可夫模型的联合似然,而是分别最大化两个隐马尔可夫模型,然后将隐马尔可夫模型融合在一起。此外,还提出了一种推理算法。通过人工验证实验对所提出的方法进行了验证。实验结果表明,与单模隐马尔可夫模型和松耦合融合模型相比,该方法显著降低了识别错误率。
In this paper, the audio and visual features of speech are integrated using a novel fused-HMM. We assume that the two sets of features may have different data rates and duration. Hidden Markov models (HMMs) are first used to model them separately, and then a general Bayesian fusion method, which is optimal in the maximum entropy sense, is employed to fuse them together. Particularly, an efficient learning algorithm is introduced. Instead of maximizing the joint likelihood of the fuse-HMM, the learning algorithm maximizes the two HMMs separately, and then fuses the HMMs together. In addition, an inference algorithm is proposed. We have tested the proposed method by person verification experiments. Results show that the proposed method significantly reduces the recognition error rates as compared to the unimodal HMMs and the loosely-coupled fusion model.