Fusing audio and visual features of speech
Fusing audio and visual features of speech
复制标题
融合语音的音频和视觉特征
DOI:
10.1109/icip.2000.899333
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Thomas S. Huang
中科院分区:
文献类型:
--
作者:
Hao Pan;Zhi;Thomas S. Huang
In this paper, the audio and visual features of speech are integrated using a novel fused-HMM. We assume that the two sets of features may have different data rates and duration. Hidden Markov models (HMMs) are first used to model them separately, and then a general Bayesian fusion method, which is optimal in the maximum entropy sense, is employed to fuse them together. Particularly, an efficient learning algorithm is introduced. Instead of maximizing the joint likelihood of the fuse-HMM, the learning algorithm maximizes the two HMMs separately, and then fuses the HMMs together. In addition, an inference algorithm is proposed. We have tested the proposed method by person verification experiments. Results show that the proposed method significantly reduces the recognition error rates as compared to the unimodal HMMs and the loosely-coupled fusion model.