Limitations of visual speech recognition

Limitations of visual speech recognition
复制标题

DOI:
--
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
Jacob L. Newman;B. Theobald;S. Cox
Jacob L. Newman;B. Theobald;S. Cox
中科院分区:
其他
文献类型:
--
作者:
Jacob L. Newman;B. Theobald;S. Cox

文献摘要

被引文献

相似文献

在这篇文章中,我们调查了自动唇读系统的局限性,我们认为如果识别器可以从其他(不可见的)语音发音装置获得额外的信息,那么可以获得的改进。隐马尔可夫模型(HMM)语音识别器使用从Mocha-Timit数据集中提取的电磁清晰度成像(EMA)数据来训练。发音信息被系统地保留在识别器中,并且性能被测试,并与典型的现有技术的唇读系统的性能进行比较。我们发现,正如预期的那样,随着发音信息的丢失,识别器的性能会下降,并且典型的唇读系统实现了类似于仅使用舌头前部信息的基于EMA的识别器的性能水平。我们的结果表明,在靠近嘴巴后部的发音器位置上有大量的信息可以利用,但即使是这样,也不足以达到声学语音识别器可以达到的相同水平的性能。
In this paper we investigate the limits of automated lip-reading systems and we consider the improvement that could be gained were additional information from other (non-visible) speech articulators available to the recogniser. Hidden Markov model (HMM) speech recognisers are trained using electromagnetic articulography (EMA) data drawn from the MOCHA-TIMIT data set. Articulatory information is systematically withheld from the recogniser and the performance is tested and compared with that of a typical state of the art lip-reading system. We find that, as expected, the performance of the recogniser degrades as articulatory information is lost, and that a typical lip-reading system achieves a level of performance similar to an EMA-based recogniser that uses information from only the front of the tongue forwards. Our results show that there is significant information in the articulator positions towards the back of the mouth that could be exploited were it available, but even this is insufficient to achieve the same level of performance as can be achieved by an acoustic speech recogniser.