课题基金 / 基金详情

Improving audio-visual speech recognition with augmented facial-mapping.

Improving audio-visual speech recognition with augmented facial-mapping.
通过增强面部映射改进视听语音识别。
批准号:
1964209
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2017
资助国家:
英国
项目状态:
已结题
起止时间:
2017 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
研究问题:是否可以通过加强新兴的面部映射技术来改进视听语音识别?实时3D面部映射和声音分区的应用能否提高视听语音识别的准确性?在撰写本文时,还没有关于使用TrueDepth摄像头的面部识别进行视听语音识别的已知研究。这可能是因为这项技术还处于起步阶段。改进的集成视听语音识别系统的潜在应用是:改善人工智能系统的人机交互。一种更便宜的自主语音治疗方法。语言学习。目标和目的本研究将专注于机器学习原理,以开发更有效的语音和面部(视觉语音)识别算法的端到端解决方案。然后,通过精确的反馈引擎,这将用于提高人类在这些领域的准确性和沟通能力。其目的是有效地结合最新的红外和邻近传感器用于实时人脸映射,以提高视听语音识别。方法由于本研究本质上是计算机科学和语言学之间的交叉学科,因此本文将首先研究当前的深度学习视听语音识别方法以及更广泛的历史语音阅读和自然语言处理技术。然后,本文将探讨苹果TrueDepth摄像头在视觉语音识别方面的潜在应用的个人准确性。TrueDepth系统主要用于面部识别和动画,基本上与微软3D Tracing Connect附件中包含的技术相同。自那以后,机器学习软件的中间件层对其进行了微型化和改进,以实现37个面部特征的实时映射和发音,精度为毫米。这项研究将首先通过记录大量的母语学习数据集并迭代有监督的深度学习算法来测试TrueDepth相机对一组视位(视觉音素)的识别精度。一旦达到可接受的视位识别准确度,这将与现有的基于音频的语音识别引擎相结合。最后阶段将评估TruDepth摄像头系统的增强是否会在统计上带来可行的改进,当与独立的语音识别引擎进行测试时。
英文摘要
Research questions:Can audio-visual speech recognition be improved through the augmentation of emerging facial mapping technology?Can the application of real-time 3D face mapping and sound compartmentalisation improve audio-visual speech recognition accuracy?Potential applications At the time of writing, no known research exists in the use of the TrueDepth camera's facial recognition for audio-visual speech recognition. This may be due to the infancy of the technology. The potential applications for an improved integrated audio-visual speech recognition system are: Improved human computer interaction for AI systems.A cheaper means of autonomous speech therapy.Language learning.Objectives and AimsThis research will focus on machine learning principles to develop a more effective end-to-end solution for speech and facial (visual speech) recognition algorithms. This will then be used to improve human accuracy and communication in these areas, through a precise feedback engine. The objective is to effectively integrate the use of the latest infrared and proximity sensors used for real-time face mapping, to improve audio-visual speech recognition.MethodologyAs this research is inherently interdisciplinary between computer science and linguistics this paper will first investigate current deep learning audio-visual speech recognition methodologies and broader historical speechreading and natural language processing techniques. This paper will then explore the individual accuracy of Apple's TrueDepth camera in terms of its potential application for visual speech recognition. The TrueDepth system is primarily used for facial recognition and animation, and is essentially the same technology contained within Microsoft's 3D tracking Connect accessory. This has since been miniaturised and improved by a middleware layer of machine learning software, to achieve the real-time mapping and articulation of 37 facial features with millimetre accurately. This research will first test the TrueDepth camera's recognition accuracy of a set visemes (visual phonemes) by recording a large native language learning dataset and iterating through a supervised deep learning algorithm. Once an acceptable level of viseme recognition accuracy is achieved, this will then be combined with an existing audio-based speech recognition engine. The final stage will assess whether the augmentation of the TruDepth camera system will result in a statistically viable improvement, when tested against standalone speech recognition engines.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金