The challenge of multispeaker lip-reading

The challenge of multispeaker lip-reading
复制标题

DOI:
--
复制
发表时间:
2008-09
期刊:
--
影响因子:
--
通讯作者:
S. Cox;R. Harvey;Yuxuan Lan;Jacob L. Newman;B. Theobald
S. Cox;R. Harvey;Yuxuan Lan;Jacob L. Newman;B. Theobald
中科院分区:
其他
文献类型:
--
作者:
S. Cox;R. Harvey;Yuxuan Lan;Jacob L. Newman;B. Theobald

文献摘要

被引文献

相似文献

在语音识别中,说话人的可变性问题已经得到了很好的研究。常见的处理方法包括对说话人的声道长度进行归一化,以及学习线性变换,将与说话人无关的模型移动到更接近新说话人的位置。在纯粹的唇读(没有音频)中,这个问题还没有得到很好的研究。结果通常是基于说话人相关的(单说话人)或多说话人(测试集中的说话人也在训练集中)数据来呈现的,这些情况在实际应用中的使用有限。这篇文章展示了在训练和测试集中不使用不同扬声器的危险。首先,我们给出了一个新的单词库AVLetters 2的分类结果,它是著名的AVLetters数据库的高清晰度版本。通过对特征的仔细选择,我们表明,对于单说话人和多说话人配置,仅视觉唇读的性能可能与仅音频识别的性能非常接近。然而,在与说话人无关的配置中,仅可视通道的性能显著下降。通过将多维尺度(MDS)应用于音频特征和视觉特征,我们证明了与通常用于音频语音识别的MFCC相比,唇读视觉特征在单个说话人内在所有口语类别中具有固有的微小变化。然而,视觉特征对说话人的身份高度敏感,而音频特征相对不变。
In speech recognition, the problem of speaker variability has been well studied. Common approaches to dealing with it include normalising for a speaker's vocal tract length and learning a linear transform that moves the speaker-independent models closer to to a new speaker. In pure lip-reading (no audio) the problem has been less well studied. Results are often presented that are based on speaker-dependent (single speaker) or multispeaker (speakers in the test-set are also in the training-set) data, situations that are of limited use in real applications. This paper shows the danger of not using different speakers in the trainingand test-sets. Firstly, we present classification results on a new single-word database AVletters 2 which is a high-definition version of the well known AVletters database. By careful choice of features, we show that it is possible for the performance of visual-only lip-reading to be very close to that of audio-only recognition for the single speaker and multi-speaker configurations. However, in the speaker independent configuration, the performance of the visual-only channel degrades dramatically. By applying multidimensional scaling (MDS) to both the audio features and visual features, we demonstrate that lip-reading visual features, when compared with the MFCCs commonly used for audio speech recognition, have inherently small variation within a single speaker across all classes spoken. However, visual features are highly sensitive to the identity of the speaker, whereas audio features are relatively invariant.