Dynamic Bayesian networks for audio-visual speech recognition

Dynamic Bayesian networks for audio-visual speech recognition
复制标题

DOI:
10.1155/s1110865702206083
复制
发表时间:
2002-11-01
期刊:
EURASIP JOURNAL ON APPLIED SIGNAL PROCESSING
影响因子:
--
通讯作者:
Murphy, K
Murphy, K
中科院分区:
其他
文献类型:
--
作者:
Nefian, AV;Liang, LH;Murphy, K

文献摘要

被引文献

相似文献

在视听语音识别(AVSR)中使用视觉特征是合理的语音生成机制,这基本上是双峰的音频和视觉表示,并需要的功能是不变的声学噪声扰动。因此,当前的AVSR系统在受声学噪声影响的环境中表现出显著的精度改进。在本文中,我们描述了使用两个统计模型的视听集成,耦合HMM(CHMM)和阶乘HMM(FHMM),并比较这些模型的性能与现有的模型用于说话人相关的视听孤立词识别。CHMM和FHMM的统计特性允许对音频和视觉观察序列的状态进行建模,同时保持它们随时间的自然相关性。在我们的实验中,CHMM整体表现最好,优于所有现有模型和FHMM。
The use of visual features in audio-visual speech recognition (AVSR) is justified by both the speech generation mechanism, which is essentially bimodal in audio and visual representation, and by the need for features that are invariant to acoustic noise perturbation. As a result, current AVSR systems demonstrate significant accuracy improvements in environments affected by acoustic noise. In this paper, we describe the use of two statistical models for audio-visual integration, the coupled HMM (CHMM) and the factorial HMM (FHMM), and compare the performance of these models with the existing models used in speaker dependent audio-visual isolated word recognition. The statistical properties of both the CHMM and FHMM allow to model the state asynchrony of the audio and visual observation sequences while preserving their natural correlation over time. In our experiments, the CHMM performs best overall, outperforming all the existing models and the FHMM.