Resolution limits on visual speech recognition

Resolution limits on visual speech recognition
复制标题

DOI:
10.1109/icip.2014.7025274
复制
发表时间:
2014-10
期刊:
2014 IEEE International Conference on Image Processing (ICIP)
影响因子:
--
通讯作者:
Helen L. Bear;R. Harvey;B. Theobald;Yuxuan Lan
Helen L. Bear;R. Harvey;B. Theobald;Yuxuan Lan
中科院分区:
其他
文献类型:
--
作者:
Helen L. Bear;R. Harvey;B. Theobald;Yuxuan Lan

文献摘要

被引文献

相似文献

只有视觉的语音识别依赖于许多难以控制的因素,例如:照明、身份、运动、情感和表情。但有些因素是可控的,比如视频分辨率,所以令人惊讶的是,目前还没有关于分辨率对唇读影响的系统研究。在这里,我们使用一个新的数据集Rosetta Raven数据来训练和测试识别器,以便我们可以测量视频分辨率对识别精度的影响。我们的结论是,与通常的做法相反,对于自动唇读来说,分辨率不必那么好。然而,当静止时下嘴唇的底部和上嘴唇的顶部之间的距离小于4个像素时,自动唇读能够可靠地工作是非常不可能的。
Visual-only speech recognition is dependent upon a number of factors that can be difficult to control, such as: lighting; identity; motion; emotion and expression. But some factors, such as video resolution are controllable, so it is surprising that there is not yet a systematic study of the effect of resolution on lip-reading. Here we use a new data set, the Rosetta Raven data, to train and test recognizers so we can measure the affect of video resolution on recognition accuracy. We conclude that, contrary to common practice, resolution need not be that great for automatic lip-reading. However it is highly unlikely that automatic lip-reading can work reliably when the distance between the bottom of the lower lip and the top of the upper lip is less than four pixels at rest.