Lipreading using convolutional neural network
Lipreading using convolutional neural network
复制标题
DOI:
10.21437/interspeech.2014-293
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
K. Noda;Yuki Yamaguchi;K. Nakadai;HIroshi G. Okuno;T. Ogata
中科院分区:
文献类型:
--
作者:
K. Noda;Yuki Yamaguchi;K. Nakadai;HIroshi G. Okuno;T. Ogata
In recent automatic speech recognition studies, deep learning architecture applications for acoustic modeling have eclipsed conventional sound features such as Mel-frequency cepstral co-efficients. However, for visual speech recognition (VSR) studies, handcrafted visual feature extraction mechanisms are still widely utilized. In this paper, we propose to apply a convolutional neural network (CNN) as a visual feature extraction mechanism for VSR. By training a CNN with images of a speaker’s mouth area in combination with phoneme labels, the CNN acquires multiple convolutional filters, used to extract visual features essential for recognizing phonemes. Further, by modeling the temporal dependencies of the generated phoneme label sequences, a hidden Markov model in our proposed sys-tem recognizes multiple isolated words. Our proposed system is evaluated on an audio-visual speech dataset comprising 300 Japanese words with six different speakers. The evaluation re-sults of our isolated word recognition experiment demonstrate that the visual features acquired by the CNN significantly out-perform those acquired by conventional dimensionality compression approaches, including principal component analysis.