ReBabel - deep learning cross-lingual vocal characteristic matching with automated lip-sync
ReBabel - deep learning cross-lingual vocal characteristic matching with automated lip-sync
批准号:
10014754
负责人:
金额:
$32.65万
依托单位:
依托单位国家:
英国
项目类别:
Collaborative R&D
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --
中文摘要
在品牌全球化、流媒体服务(Netflix/Amazon Prime)、社交平台(YouTube/Twitch)和智能手机无处不在的推动下,外语制作的需求不断增长,分别为21亿GB/年(2019年)和5.9%的CAGR(Marketwatch.com)。欧洲约占全球配音业的22%。然而,尽管观众数据和调查都显示,相对于字幕,观众更喜欢配音内容,但跨语言配音的使用受到成本高和制作时间长的限制。其他挑战包括:*改编剧本的难度,以及需要多个镜头才能让配音演员(VAS)对口型进行匹配。糟糕的假唱不吸引人,有时还很滑稽。*选角。理想的情况是,寻找与原演员非常匹配的演员,但受时间和成本的限制,这甚至可能导致最佳匹配的方法和/或重复使用VAS。*声音/身体不匹配导致的审美不连贯。*与多个明星相关的VAS的时间瓶颈。因此,制片人必须权衡配音的额外成本和增加销售的可能性,而不是只提供字幕。**对项目的愿景**在这个项目中,我们将基于我们的语音合成/深度学习(DL)专业知识来创建ReBabel,这是一种创新且独特的跨语言配音工具,可以降低成本,通过关键创新提高质量并应对挑战:1.跨语言语音变形:将外语表演改变为听起来像是原始表演者给出的声音。我们独特的技术通过利用我们独特的神经网络模型实现语音到语音的合成(在不同的语言中将一种声音转换为另一种声音),该模型将语音转换为概念并返回。我们的概念抽象既传达了字面上的文本意义/信息,也传达了副语言学(有时称为发音)信息,这些信息可以修改意义、赋予细微差别和传达情感,并可以编辑以应用所需的转换。音频-口型同步:人工智能驱动的手动/自动音频重新同步,以匹配视频中的嘴唇运动。这将包括直接修改音频通道以匹配演员的嘴唇,同时传达预期的声音信息/内容/表演。ReBabel将在我们正在开发的Voice Studio套件中形成一个模块,该模块最初将在最终的消费者/消费者版本之前向专业制作公司提供“Photoshop for Voice”。**主要目标**开发快速、准确和高质量的机器学习(ML)模型和支持数据结构。*为关键的国际语言实施DL,以实现语音/视频唇形同步。*测试和验证模型结果。*为用户的技术知识(可用性)实现较低的标准。
英文摘要
At £2.1bn/annum (2019) and 5.9% CAGR (Marketwatch.com) there is growing demand for foreign-language production, driven by brand globalisation, streaming services (Netflix/Amazon Prime), social platforms (YouTube/Twitch) and the ubiquity of smart phones. Europe accounts for ~22% of the global dubbing industry.However, while audience data and surveys both show a preference for dubbed content over subtitles, the use of cross-lingual dubbing is limited by high costs and long production times. Other challenges include:* The difficulty of adapting scripts and the need for multiple takes so Voice Actors (VAs) can match lip-synch. Poor lip-synch is unappealing and sometimes comical.* Casting. Ideally a close match to the original actor is sought but is constrained by time and costs that may even result in a "best-fit" approach and/or reusing VAs.* Aesthetic incoherence arising from a voice/body mismatch.* Time bottlenecks when VAs associated with multiple stars have limited availability.Producers must therefore balance the additional cost of a dub and the potential for increased sales against offering subtitles only.**Vision for the project**In this project we will build on our Speech-Synthesis/Deep Learning (DL) expertise to create ReBabel, an innovative and unique new cross-lingual dubbing tool that lowers costs, improves quality and addresses the challenges via key innovations:1. Cross-lingual voice-morphing: Changing a foreign-language performance to sound like it was given by the original performer. Our unique technology enables speech-to-speech synthesis (transforming one voice to another in a different language) by utilising our unique neural network model that transforms speech into concepts and back. Our conceptual abstractions convey both literal textual meaning/information, as well as the paralinguistics (sometimes called vocalics) information that can modify meaning, give nuance, and convey emotion and can be edited to apply the desired transformations.2. Audio-LipSync: AI-driven manual/automatic re-synchronising of the audio to match the lip movements in the video. This will involve direct modification of the audio channel to match the lips of the actor, while conveying the vocal message/content/performance that is expected.ReBabel will form a module within our in-development Voice Studio suite that will initially offer "Photoshop for Voice" to professional production companies before eventual consumer/prosumer versions.**Key objectives*** Development of a fast, accurate and high-quality Machine Learning (ML) model and supporting data structures.* Implementation of DL for key international languages to enable voice/video lip-sync.* Testing and validation of model results.* Achieving a low bar for user's technical knowledge (usability).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于深穿透拉曼光谱的安全光照剂量的深层病灶无创检测与深度预测
-
批准号:82372016
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:林俐
-
依托单位:
GREB1突变介导雌激素受体信号通路导致深部浸润型子宫内膜异位症的分子遗传机制研究
-
批准号:82371652
-
项目类别:面上项目
-
资助金额:45.00万元
-
批准年份:2023
-
负责人:刘开江
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位: