Speaker-independent machine lip-reading with speaker-dependent viseme classifiers

Speaker-independent machine lip-reading with speaker-dependent viseme classifiers
复制标题

DOI:
--
复制
发表时间:
2015-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Helen L. Bear;S. Cox;R. Harvey
Helen L. Bear;S. Cox;R. Harvey
中科院分区:
其他
文献类型:
--
作者:
Helen L. Bear;S. Cox;R. Harvey

文献摘要

被引文献

相似文献

在机器唇读中,这是从视觉信息中识别语音,有证据表明视觉语音高度依赖于说话者[1]。在这里,我们使用音素聚类方法来形成新的音素到视位映射为个人和多个扬声器。我们使用这些地图来研究说话者在视觉上的相似程度。我们的结论是,广义上讲,扬声器有相同的剧目的嘴的姿态,他们不同的是在使用的姿态。
In machine lip-reading, which is identification of speech from visual-only information, there is evidence to show that visual speech is highly dependent upon the speaker [1]. Here, we use a phoneme-clustering method to form new phoneme-to-viseme maps for both individual and multiple speakers. We use these maps to examine how similarly speakers talk visually. We conclude that broadly speaking, speakers have the same repertoire of mouth gestures, where they differ is in the use of the gestures.