Lipreading using convolutional neural network

Lipreading using convolutional neural network
复制标题

DOI:
10.21437/interspeech.2014-293
复制
发表时间:
2014
期刊:
--
影响因子:
--
通讯作者:
K. Noda;Yuki Yamaguchi;K. Nakadai;HIroshi G. Okuno;T. Ogata
K. Noda;Yuki Yamaguchi;K. Nakadai;HIroshi G. Okuno;T. Ogata
中科院分区:
其他
文献类型:
--
作者:
K. Noda;Yuki Yamaguchi;K. Nakadai;HIroshi G. Okuno;T. Ogata

文献摘要

相似文献

在最近的自动语音识别研究中,用于声学建模的深度学习架构应用已经超过了传统的声音特征,如Mel频率倒谱系数。然而,对于视觉语音识别(VSR)的研究,手工制作的视觉特征提取机制仍然被广泛使用。在本文中,我们提出将卷积神经网络(CNN)作为VSR的视觉特征提取机制。通过使用说话者嘴部区域的图像与音素标签相结合来训练CNN,CNN获得了多个卷积滤波器,用于提取识别音素所必需的视觉特征。此外,通过对生成的音素标签序列的时间依赖性建模,隐马尔可夫模型在我们提出的系统中识别多个孤立的单词。我们提出的系统进行评估的视听语音数据集,包括300个日语单词与六个不同的扬声器。我们孤立词识别实验的评估结果表明,CNN获得的视觉特征显著优于传统的维度压缩方法,包括主成分分析。
In recent automatic speech recognition studies, deep learning architecture applications for acoustic modeling have eclipsed conventional sound features such as Mel-frequency cepstral co-efficients. However, for visual speech recognition (VSR) studies, handcrafted visual feature extraction mechanisms are still widely utilized. In this paper, we propose to apply a convolutional neural network (CNN) as a visual feature extraction mechanism for VSR. By training a CNN with images of a speaker’s mouth area in combination with phoneme labels, the CNN acquires multiple convolutional filters, used to extract visual features essential for recognizing phonemes. Further, by modeling the temporal dependencies of the generated phoneme label sequences, a hidden Markov model in our proposed sys-tem recognizes multiple isolated words. Our proposed system is evaluated on an audio-visual speech dataset comprising 300 Japanese words with six different speakers. The evaluation re-sults of our isolated word recognition experiment demonstrate that the visual features acquired by the CNN significantly out-perform those acquired by conventional dimensionality compression approaches, including principal component analysis.