课题基金 / 基金详情

Highly expressive voices for machine video content localisation

Highly expressive voices for machine video content localisation
用于机器视频内容本地化的高表现力声音
批准号:
73674
负责人:
金额:
$43.61万
依托单位国家:
英国
项目类别:
Study
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
想象一下,任何语言的视频都可以使用,既有原始演员独特的声音品质,也有他们用新语言保留下来的独特的台词表达方式。这是Papercup将利用机器学习世界的最新发展来实现的雄心勃勃的愿景。2016年,谷歌的DeepMind创造了WaveNet声码器。这是语音合成领域的一场革命。在此之前,语音合成模型要么是串联的(这意味着它们通过将录制的语音的短音频样本粘合在一起来工作),要么是模型化的方法,即使用人类语音产生系统如何工作的模型从零开始生成语音。串联合成通常会产生更自然的声音,但会产生不自然的流动,因为音频样本来自不相关的语音部分。模型化的方法往往能产生更好的流畅性,但声音听起来像是机器人。WaveNET是一种深度学习方法,直接在音频样本上进行训练,并结合了建模方法的自然变化和拼接方法的自然声音。这一发展意味着语音合成可能变得与人类的语音本质上无法区分。然而,声码器(如WaveNet)甚至还不是故事的一半。你仍然要告诉它该说什么,怎么说。要让计算机做到这一点,我们必须首先识别原始视频中说了什么,是谁说的,是以什么方式说的。Papercup利用了深度学习的最新发展,并开发了一种正在申请专利的方法,用于分析每个扬声器的独特声学特征,以及他们说话的方式。这是由我们的算法使用内部学习表示进行编码的,这使得重音、语调和情感能够跨语言传递,类似于翻译工具将文本从一种语言翻译成另一种语言的方式。通过这种方式,Papercup的方法复制了演员独特的发声特征,并复制了他们的表达。这有可能给画外音翻译行业带来革命性的变化,因为它创建了忠实的画外音翻译,可以用其他语言准确地传达原始内容,而且与使用传统的画外音翻译服务相比,成本要低得多。
英文摘要
Imagine any video available in any language, with both the unique qualities of the original actors' voices, and the unique way in which they delivered their lines, preserved in the new language. This is the ambitious vision that Papercup will make reality by harnessing the latest developments in the world of machine learning.In 2016, Google's Deepmind created the WaveNet vocoder. This was a revolution in speech synthesis. Prior to this, speech synthesis models were either concatenative (meaning that they work by glueing together short audio samples of recorded speech) or modelled methods, which generate speech "from scratch" using a model of how the human speech production system works. Concatenative synthesis typically resulted in more natural sounding voices, but with unnatural flow because the audio samples come from unrelated sections of speech. Modelled methods tended to produce better flow, but the voices sounded robotic. WaveNet is a deep-learning method, trained directly on audio samples, and combines the natural variation of modelled methods with the natural sound of concatenative methods. This development means that speech synthesis could become essentially indistinguishable from human speech.A vocoder (such as WaveNet), however, is not even half the story. You still have to tell it what to say, and how to say it. For a computer to achieve this, we must first recognise what was said in the original video, by whom, and in what way. Papercup exploits the latest developments in deep learning and has developed a patent-pending method for analysing the unique acoustic features of each speaker, and the way in which they delivered their lines. This is encoded by our algorithms using an internal learned representation, which enables the stresses, intonation, and emotion to be transferred across languages, in a manner analogous to the way translation tools translate text from one language to another.In this way, Papercup's approach replicates the unique vocal characteristics of the actors, and replicates their delivery. This has the potential to revolutionise the voiceover translation industry by creating faithful voiceover translations that accurately convey the original content in additional languages, and do so at scale with significantly lower costs than using traditional voiceover translation services with voice-actors.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
脑梗塞运动性失语后语言功能恢复机制的fMRI功能连接研究
  • 批准号:
    30700193
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    18.0万元
  • 批准年份:
    2007
  • 负责人:
    张权
  • 依托单位: