ReBabel - deep learning cross-lingual vocal characteristic matching with automated lip-sync
ReBabel - deep learning cross-lingual vocal characteristic matching with automated lip-sync
批准号:
10014754
负责人:
金额:
$32.65万
依托单位:
依托单位国家:
英国
项目类别:
Collaborative R&D
财政年份:
2022
资助国家:
英国
项目状态:
已结题
起止时间:
2022 至 --
中文摘要
在21亿英镑/年(2019年)和5.9%的复合年增长率(Marketwatch.com)下,对外语制作的需求不断增长,这是由品牌全球化,流媒体服务(Netflix/亚马逊Prime),社交平台(YouTube/Twitch)和智能手机的普及所推动的。欧洲约占全球配音行业的22%。然而,尽管观众数据和调查都显示配音内容比字幕更受欢迎,但跨语言配音的使用受到成本高和制作时间长的限制。其他挑战包括:* 改编脚本的难度和需要多次拍摄,以便配音演员(VA)可以匹配对口型。糟糕的对口型很不吸引人,有时还很滑稽。铸造理想情况下,寻求与原始参与者的密切匹配,但受到时间和成本的限制,甚至可能导致“最适合”的方法和/或重复使用VA。声音/身体不匹配导致的美学不连贯。当与多位明星相关的VA的可用性有限时,时间瓶颈。因此,制片人必须平衡配音的额外成本和增加销售的潜力,而不是只提供字幕。**项目愿景 ** 在此项目中,我们将利用语音合成/深度学习(DL)专业知识,创建ReBabel,这是一种创新且独特的跨语言配音工具,可通过关键创新降低成本,提高质量并应对挑战:1.跨语言变声:改变外语表演,使其听起来像是原表演者的声音。我们独特的技术通过利用我们独特的神经网络模型将语音转换为概念并返回,从而实现语音到语音合成(将一种语音转换为另一种不同语言的语音)。我们的概念抽象既传达了文字的文本意义/信息,也传达了可以修改意义、赋予细微差别和传达情感的非语言学(有时称为声乐)信息,并且可以进行编辑以应用所需的转换。Audio-LipSync:人工智能驱动的手动/自动音频重新同步,以匹配视频中的嘴唇运动。这将涉及直接修改音频通道以匹配演员的嘴唇,同时传达预期的声音信息/内容/表演。ReBabel将在我们正在开发的Voice Studio套件中形成一个模块,最初将在最终的消费者/专业消费者版本之前向专业制作公司提供“Photoshop for Voice”。**关键目标 * 开发快速、准确和高质量的机器学习(ML)模型和支持数据结构。针对主要国际语言实施DL,以实现语音/视频口型同步。*模型结果的测试和验证。实现用户技术知识(可用性)的低标准。
英文摘要
At £2.1bn/annum (2019) and 5.9% CAGR (Marketwatch.com) there is growing demand for foreign-language production, driven by brand globalisation, streaming services (Netflix/Amazon Prime), social platforms (YouTube/Twitch) and the ubiquity of smart phones. Europe accounts for ~22% of the global dubbing industry.However, while audience data and surveys both show a preference for dubbed content over subtitles, the use of cross-lingual dubbing is limited by high costs and long production times. Other challenges include:* The difficulty of adapting scripts and the need for multiple takes so Voice Actors (VAs) can match lip-synch. Poor lip-synch is unappealing and sometimes comical.* Casting. Ideally a close match to the original actor is sought but is constrained by time and costs that may even result in a "best-fit" approach and/or reusing VAs.* Aesthetic incoherence arising from a voice/body mismatch.* Time bottlenecks when VAs associated with multiple stars have limited availability.Producers must therefore balance the additional cost of a dub and the potential for increased sales against offering subtitles only.**Vision for the project**In this project we will build on our Speech-Synthesis/Deep Learning (DL) expertise to create ReBabel, an innovative and unique new cross-lingual dubbing tool that lowers costs, improves quality and addresses the challenges via key innovations:1. Cross-lingual voice-morphing: Changing a foreign-language performance to sound like it was given by the original performer. Our unique technology enables speech-to-speech synthesis (transforming one voice to another in a different language) by utilising our unique neural network model that transforms speech into concepts and back. Our conceptual abstractions convey both literal textual meaning/information, as well as the paralinguistics (sometimes called vocalics) information that can modify meaning, give nuance, and convey emotion and can be edited to apply the desired transformations.2. Audio-LipSync: AI-driven manual/automatic re-synchronising of the audio to match the lip movements in the video. This will involve direct modification of the audio channel to match the lips of the actor, while conveying the vocal message/content/performance that is expected.ReBabel will form a module within our in-development Voice Studio suite that will initially offer "Photoshop for Voice" to professional production companies before eventual consumer/prosumer versions.**Key objectives*** Development of a fast, accurate and high-quality Machine Learning (ML) model and supporting data structures.* Implementation of DL for key international languages to enable voice/video lip-sync.* Testing and validation of model results.* Achieving a low bar for user's technical knowledge (usability).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于深穿透拉曼光谱的安全光照剂量的深层病灶无创检测与深度预测
-
批准号:82372016
-
项目类别:面上项目
-
资助金额:48.00万元
-
批准年份:2023
-
负责人:林俐
-
依托单位:
GREB1突变介导雌激素受体信号通路导致深部浸润型子宫内膜异位症的分子遗传机制研究
-
批准号:82371652
-
项目类别:面上项目
-
资助金额:45.00万元
-
批准年份:2023
-
负责人:刘开江
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位: