Lip Location Normalized Training for Visual Speech Recognition

Lip Location Normalized Training for Visual Speech Recognition
复制标题

视觉语音识别的嘴唇位置标准化训练

DOI:
--
复制
发表时间:
2000
影响因子:
0.7
通讯作者:
T. Kitamura
T. Kitamura
中科院分区:
计算机科学4区
文献类型:
--
作者:
O. Vanegas;K. Tokuda;T. Kitamura

文献摘要

被引文献

相似文献

为了提高基于视觉信息的语音识别系统的性能,本文提出了一种唇形位置归一化的方法。基本上,有两种类型的信息在语音识别过程中有用;第一个是语音信号本身,第二个是运动中的嘴唇的视觉信息。本文试图解决使用运动中的嘴唇图像所带来的一些问题,如嘴唇位置变化所产生的效果。本文提出的唇部位置归一化方法是基于一种唇部位置搜索算法,该算法将位置归一化融入到模型训练中。在Tulips1和M2VTS数据库上进行了独立于说话人的孤立词识别实验。实验表明,该方法对M2VTS数据库的十位数词识别识别率为74.5%,错误率为35.7%。关键词:隐马尔可夫模型,唇定位归一化,唇读,Tulips1, M2VTS
This paper describes a method to normalize the lip position for improving the performance of a visualinformation-based speech recognition system. Basically, there are two types of information useful in speech recognition processes; the first one is the speech signal itself and the second one is the visual information from the lips in motion. This paper tries to solve some problems caused by using images from the lips in motion such as the effect produced by the variation of the lip location. The proposed lip location normalization method is based on a search algorithm of the lip position in which the location normalization is integrated into the model training. Experiments of speaker-independent isolated word recognition were carried out on the Tulips1 and M2VTS databases. Experiments showed a recognition rate of 74.5% and an error reduction rate of 35.7% for the ten digits word recognition M2VTS database. key words: hidden Markov model, lip location normalization, lipreading, Tulips1, M2VTS