Audio-visual speech recognition using deep bottleneck features and high-performance lipreading

Audio-visual speech recognition using deep bottleneck features and high-performance lipreading
复制标题

DOI:
10.1109/apsipa.2015.7415335
复制
发表时间:
2015-12
期刊:
2015 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)
影响因子:
--
通讯作者:
S. Tamura;H. Ninomiya;N. Kitaoka;Shin Osuga;Y. Iribe;K. Takeda;S. Hayamizu
S. Tamura;H. Ninomiya;N. Kitaoka;Shin Osuga;Y. Iribe;K. Takeda;S. Hayamizu
中科院分区:
其他
文献类型:
--
作者:
S. Tamura;H. Ninomiya;N. Kitaoka;Shin Osuga;Y. Iribe;K. Takeda;S. Hayamizu

文献摘要

相似文献

本文通过(1)探索高性能的视觉特征,(2)应用音频和视觉深度瓶颈特征来提高AVSR性能,以及(3)研究视觉模式下语音活动检测的有效性,开发了一种视听语音识别(AVSR)方法。在我们的方法中,纳入了多种视觉特征,然后通过深度学习技术将其转化为瓶颈特征。利用所提出的特征,我们成功地在与说话人无关的开放条件下实现了73.66%的读唇精度,在噪声环境下平均达到了90%左右的AVSR精度。此外,我们从视觉特征中提取语音片段,得到了77.80%的读唇准确率。发现VAD在音频和视觉两种模式下都很有用,可以更好地读唇和AVSR。
This paper develops an Audio-Visual Speech Recognition (AVSR) method, by (1) exploring high-performance visual features, (2) applying audio and visual deep bottleneck features to improve AVSR performance, and (3) investigating effectiveness of voice activity detection in a visual modality. In our approach, many kinds of visual features are incorporated, subsequently converted into bottleneck features by deep learning technology. By using proposed features, we successfully achieved 73.66% lipreading accuracy in speaker-independent open condition, and about 90% AVSR accuracy on average in noisy environments. In addition, we extracted speech segments from visual features, resulting 77.80% lipreading accuracy. It is found VAD is useful in both audio and visual modalities, for better lipreading and AVSR.