课题基金 / 基金详情

Improving audio-visual speech recognition with augmented facial-mapping.

Improving audio-visual speech recognition with augmented facial-mapping.
通过增强面部映射改进视听语音识别。
批准号:
1964209
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Research questions:Can audio-visual speech recognition be improved through the augmentation of emerging facial mapping technology?Can the application of real-time 3D face mapping and sound compartmentalisation improve audio-visual speech recognition accuracy?Potential applications At the time of writing, no known research exists in the use of the TrueDepth camera's facial recognition for audio-visual speech recognition. This may be due to the infancy of the technology. The potential applications for an improved integrated audio-visual speech recognition system are: Improved human computer interaction for AI systems.A cheaper means of autonomous speech therapy.Language learning.Objectives and AimsThis research will focus on machine learning principles to develop a more effective end-to-end solution for speech and facial (visual speech) recognition algorithms. This will then be used to improve human accuracy and communication in these areas, through a precise feedback engine. The objective is to effectively integrate the use of the latest infrared and proximity sensors used for real-time face mapping, to improve audio-visual speech recognition.MethodologyAs this research is inherently interdisciplinary between computer science and linguistics this paper will first investigate current deep learning audio-visual speech recognition methodologies and broader historical speechreading and natural language processing techniques. This paper will then explore the individual accuracy of Apple's TrueDepth camera in terms of its potential application for visual speech recognition. The TrueDepth system is primarily used for facial recognition and animation, and is essentially the same technology contained within Microsoft's 3D tracking Connect accessory. This has since been miniaturised and improved by a middleware layer of machine learning software, to achieve the real-time mapping and articulation of 37 facial features with millimetre accurately. This research will first test the TrueDepth camera's recognition accuracy of a set visemes (visual phonemes) by recording a large native language learning dataset and iterating through a supervised deep learning algorithm. Once an acceptable level of viseme recognition accuracy is achieved, this will then be combined with an existing audio-based speech recognition engine. The final stage will assess whether the augmentation of the TruDepth camera system will result in a statistically viable improvement, when tested against standalone speech recognition engines.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金