Visual speech recognition: a solution from feature extraction to words classification
Visual speech recognition: a solution from feature extraction to words classification
复制标题
视觉语音识别:从特征提取到单词分类的解决方案
DOI:
10.1109/sibgra.2003.1241036
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
D. Borges
中科院分区:
文献类型:
--
作者:
Luciana Gonçalves da Silveira;J. Facon;D. Borges
Audio-visual speech recognition has been an active area of research lately. A bit, and yet unsolved part of this problem is the visual only recognition, or lip reading. Considering an image sequence of a person pronouncing a word, a full image analysis solution would have to segment the mouth area, extract relevant features, and use them to be able to classify the word from those visual features. We approach this problem by proposing a segmentation technique for the lips contours together with a set of features based on the extracted contours which is able to perform lip reading with promising results. We have collected visual speech sequences in our lab and show the results for a set of ten words in Brazilian Portuguese, spoken by different speakers in more than 150 samples. The approach can be extended and applied to other spoken languages as well.