Visual speech recognition: a solution from feature extraction to words classification

Visual speech recognition: a solution from feature extraction to words classification
复制标题

视觉语音识别:从特征提取到单词分类的解决方案

DOI:
10.1109/sibgra.2003.1241036
复制
发表时间:
2003
期刊:
16th Brazilian Symposium on Computer Graphics and Image Processing (SIBGRAPI 2003)
影响因子:
--
通讯作者:
D. Borges
D. Borges
中科院分区:
--
文献类型:
--
作者:
Luciana Gonçalves da Silveira;J. Facon;D. Borges

文献摘要

被引文献

相似文献

视听语音识别最近一直是一个活跃的研究领域。这个问题中尚未解决的部分是仅视觉识别或唇读。考虑到一个人发音的图像序列,完整的图像分析解决方案必须分割嘴部区域,提取相关特征,并使用它们来根据这些视觉特征对单词进行分类。我们通过提出一种嘴唇轮廓分割技术以及基于提取轮廓的一组特征来解决这个问题,该技术能够执行唇读并取得有希望的结果。我们在实验室中收集了视觉语音序列,并显示了 150 多个样本中不同说话者所说的一组 10 个巴西葡萄牙语单词的结果。该方法也可以扩展并应用于其他口语。
Audio-visual speech recognition has been an active area of research lately. A bit, and yet unsolved part of this problem is the visual only recognition, or lip reading. Considering an image sequence of a person pronouncing a word, a full image analysis solution would have to segment the mouth area, extract relevant features, and use them to be able to classify the word from those visual features. We approach this problem by proposing a segmentation technique for the lips contours together with a set of features based on the extracted contours which is able to perform lip reading with promising results. We have collected visual speech sequences in our lab and show the results for a set of ten words in Brazilian Portuguese, spoken by different speakers in more than 150 samples. The approach can be extended and applied to other spoken languages as well.