Extraction of visual features for lipreading

Extraction of visual features for lipreading
复制标题

DOI:
10.1109/34.982900
复制
发表时间:
2002-02-01
影响因子:
23.6
通讯作者:
Harvey, R
Harvey, R
中科院分区:
计算机科学1区
文献类型:
--
作者:
Matthews, I;Cootes, TF;Harvey, R

文献摘要

被引文献

相似文献

语音的多模态特性在人机交互中经常被忽略,但是嘴唇变形和其他身体运动,例如头部的运动,传达了额外的信息。我们整合了来自许多来源的语音提示,这提高了可懂度,特别是当声学信号降级时。本文展示了如何额外的,往往是互补的,视觉语音信息可以用于语音识别。三种方法的参数化唇图像序列识别使用隐马尔可夫模型进行了比较。其中两个是自上而下的方法,适合的内部和外部唇轮廓的模型,并从形状或形状和外观的主成分分析,分别获得唇读功能。第三,自下而上,方法使用非线性尺度空间分析直接从像素强度形成特征。所有的方法进行了比较,一个多人的视觉语音识别任务的孤立的字母。
The multimodal nature of speech is often ignored in human-computer interaction, but lip deformations and other body motion, such as those of the head, convey additional information. We integrate speech cues from many sources and this improves intelligibility, especially when the acoustic signal is degraded. This paper shows how this additional, often complementary, visual speech information can be used for speech recognition. Three methods for parameterizing lip image sequences for recognition using hidden Markov models are compared. Two of these are top-down approaches that fit a model of the inner and outer lip contours and derive lipreading features from a principal component analysis of shape or shape and appearance, respectively. The third, bottom-up, method uses a nonlinear scale-space analysis to form features directly from the pixel intensity. All methods are compared on a multitalker visual speech recognition task of isolated letters.