Bimodal speech recognition using lip movement measured by optical flow analysis

Bimodal speech recognition using lip movement measured by optical flow analysis
复制标题

使用通过光流分析测量的嘴唇运动进行双模语音识别

DOI:
--
复制
发表时间:
2001
期刊:
--
影响因子:
--
通讯作者:
K. Iwano
K. Iwano
中科院分区:
--
文献类型:
--
作者:
K. Iwano

文献摘要

被引文献

相似文献

本文提出了一种利用光流分析测量嘴唇运动的双模语音识别方法。光流被定义为亮度图案运动的视速度的分布。由于可以在不提取说话人嘴唇轮廓和位置的情况下计算光流,因此可以获得关于嘴唇运动的稳健的视觉特征。我们的方法在每一帧中计算两个视觉特征:流速的水平分量和垂直分量的方差。由于这些特征表示说话人嘴巴的运动,因此它们对于估计受噪声污染的语音中的停顿/静默周期特别有用。在基于隐马尔可夫模型的识别框架中,将视觉特征和声学特征相结合。使用从干净语音数据中提取的组合特征来训练音素HMM。在识别受噪声污染的语音时,对视觉特征的观察概率进行加权。使用11名男性说话者发出相连数字的视听数据进行了实验。通过仅针对静音HMM合并视觉信息,单词准确率比纯音频识别方案实现了以下改进;在SNR=5db时为5%,在SNR=10db时为12%。
This paper proposes bimodal speech recognition using lip movements measured by optical-flow analysis. The optical flow is defined as the distribution of apparent velocities of brightness pattern movements. Since the optical flow can be computed without extracting the speaker’s lip contours and location, robust visual features can be obtained on lip movements. Our method calculates two visual features in each frame: variances of horizontal and vertical components of flow velocities. Since these features represent movement of the speaker’s mouth, they are especially useful for estimating pause/silence periods in noise-corrupted speech. The visual features are combined with acoustic features in the framework of HMM-based recognition. Phoneme HMMs are trained using the combined features extracted from clean speech data. In recognizing noise-corrupted speech, the observation probability of visual features are weighted. Experiments have been carried out using audio-visual data by 11 male speakers uttering connected digits. The following improvements of word accuracy over the audio-only recognition scheme were achieved by combining visual information only for silence HMM; 5% at SNR=5dB and 12% at SNR=10dB.