Large-vocabulary audio-visual speech recognition by machines and humans

Large-vocabulary audio-visual speech recognition by machines and humans
复制标题

机器和人类的大词汇量视听语音识别

DOI:
--
复制
发表时间:
2001
期刊:
Interspeech
影响因子:
--
通讯作者:
Eric Helmuth
Eric Helmuth
中科院分区:
--
文献类型:
--
作者:
G. Potamianos;C. Neti;G. Iyengar;Eric Helmuth

文献摘要

被引文献

相似文献

我们比较自动识别与人类感知的视听语音,在大词汇量,连续语音识别(LVCSR)域。具体来说,我们研究了机器和人类的视觉模态的好处,当结合音频退化的语音串音噪声在各种信噪比(SNR)。我们首先考虑一个自动speechreading系统与基于像素的视觉前端,使用功能融合的双峰集成,我们比较其性能与音频的LVCSR系统。然后,我们描述了人类的语音感知实验的结果,其中受试者被要求转录音频和视听话语在不同的信噪比。对于机器和人类,我们观察到与10 dB的仅音频性能相比约6 dB的有效SNR增益,然而这种增益在其他SNR下显著偏离。此外,自动视听识别在低SNR下优于人类仅音频语音感知。
We compare automatic recognition with human perception of audio-visual speech, in the large-vocabulary, continuous speech recognition (LVCSR) domain. Specifically, we study the benefit of the visual modality for both machines and humans, when combined with audio degraded by speech-babble noise at various signal-to-noise ratios (SNRs). We first consider an automatic speechreading system with a pixel based visual front end that uses feature fusion for bimodal integration, and we compare its performance with an audio-only LVCSR system. We then describe results of human speech perception experiments, where subjects are asked to transcribe audio-only and audiovisual utterances at various SNRs. For both machines and humans, we observe approximately a 6 dB effective SNR gain compared to the audio-only performance at 10 dB, however such gains significantly diverge at other SNRs. Furthermore, automatic audio-visual recognition outperforms human audioonly speech perception at low SNRs.