Exemplar-based Expressive Speech Synthesis
Exemplar-based Expressive Speech Synthesis
批准号:
EP/V046772/1
负责人:
Anton Ragni
金额:
$27.81万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2021
资助国家:
英国
项目状态:
已结题
起止时间:
2021 至 --
中文摘要
合成语音正变得无处不在:家里的智能扬声器,公共交通上的广播系统,以及呼叫线上的语音助手。公众对“更聪明”的助手有着强烈的需求,他们能够对我们的笑话大笑;能够作为鼓励和强调的导师与我们的孩子互动;打电话来检查我们的父母;为孤独的人提供令人放心的“耳朵”;以及提供平静和支持性的虚拟治疗。为了支持当前和未来的应用,语音合成技术需要满足许多要求。首先,它需要可定制,以便快速研发,其次,它需要能够产生任何语音内容,包括富有表现力的声音特征。然而,目前的合成技术都不能同时满足上述所有要求。例如,虽然目前的非机器学习方法允许将预先记录的短语有效地组合成完整的句子,但这也意味着必须首先记录缺失的必要短语,从而限制了它们的灵活性和效率。另一方面,目前的机器学习模型可以无缝地合成任何口语内容。然而,创建这样的模型是一个非常昂贵、耗时和计算要求高的过程。此外,这些模型对语音特征的质量提供的控制非常有限,并且缺乏可解释性,这在研究和商业环境中都是非常理想的条件。在这个项目中,目标是通过借鉴认知科学中的样本的概念来开发一种计算高效、可定制、可表达和可解释的语音合成。在认知科学领域,样本和原型的概念形成了关于人类如何对概念进行分类的重要观点的一部分。特别是,范例理论认为,单一的例子,而不是原型(平均的例子),形成了我们如何理解世界和与世界互动的基本构件。支持范例理论的关键论点是,作为人类,我们有能力仅基于几个例子来解决复杂的任务,这使得这个理论对涉及复杂现象或需要高计算效率的应用程序很有吸引力。此外,表现性语音合成结合了表现力和语音产生,这是两个仍然知之甚少的复杂现象。与原型理论不同,样例理论至少在理论上能够产生富有表现力的语音,前提是至少有一个期望的口头内容的记录和一个具有期望的表现力的记录是可用的。最后,采用范例理论通过使用可以在任务中检查、修改、替换、添加等的真实实例来提高决策过程的透明度。这一目标将通过三个创新手段来实现:i)建立基于范例的语音合成的方法学框架;ii)从预先记录的范例建立基于范例的语音表达;iii)提出一种新的方法来将这种基于表现力的表征整合到i)的框架中。
英文摘要
Synthetic voices are becoming ubiquitous: `smart' speakers at home, announcement systems on public transport, and voice-enabled assistants on call lines. There exist a strong public demand for `smarter' assistants capable of laughing at our jokes; interacting with our children as encouraging and emphatic tutors; calling to check up on our parents; providing a reassuring `ear' for an isolated person; and offering calming and supportive virtual therapy. To support current and future applications, voice synthesis technology needs to satisfy a number of requirements. First, it needs to be customisable for rapid research and development, and second, it needs to be able to produce any spoken content, including expressive voice characteristics. However, none of the current synthesis technologies can simultaneously satisfy all of the above requirements. For instance, while current non-machine learning approaches allow pre-recorded phrases to be efficiently combined into complete sentences, it also means that missing necessary phrases must be recorded first, thereby limiting their flexibility and efficiency. On the other hand, current machine learning models can seamlessly synthesise any spoken content. However, creating such models is a very costly, time-consuming and computationally demanding process. Furthermore, these models offer a very limited control over the qualities of the voice characteristics and lack interpretability, which are highly desirable conditions in both research and commercial settings.In this project, the objective is to develop a computationally efficient, customisable, expressive and interpretable speech synthesis, by drawing from the concept of `exemplars' in cognitive science.In the field of cognitive science, the notions of `exemplars' and `prototypes' form a part of a prominent view on how humans categorise concepts. In particular, exemplar theory argues that singular examples, rather than prototypes (an average of examples), form the basic building blocks of how we understand and interact with the world. The key argument in favour of exemplar theory is our ability as humans to solve complex tasks based on just a few examples, which makes this theory appealing to applications that involve complex phenomena or that require high computational efficiency. Furthermore, expressive speech synthesis combines expressivity and speech production, which are two complex phenomena that remain poorly understood. Unlike prototype theory, exemplar theory, at least theoretically, enables to produce expressive speech, provided that at least one recording of the desired spoken content and one recording featuring the desired expressivity are available. Lastly, adopting exemplar theory promotes transparency during the decision making process through the use of real examples that can be inspected, modified, replaced, added, etc. within the task.The objective will be achieved through three innovative means by: i) formulating a methodological framework for exemplar-based speech synthesis, ii) building an exemplar-based representation for speech expressivity from pre-recorded examples and iii) presenting a novel methodology for integrating this expressivity-based representation into the framework of i).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
-
批准号:--
-
项目类别:外国青年学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:江洋子
-
依托单位:
Incentive and governance schenism study of corporate green washing behavior in China: Based on an integiated view of econfiguration of environmental authority and decoupling logic
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:YU BYUNGJUN
-
依托单位:
Exploring the Intrinsic Mechanisms of CEO Turnover and Market Reaction: An Explanation Based on Information Asymmetry
-
批准号:W2433169
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:HAOFEI ZHANG
-
依托单位:
含Re、Ru先进镍基单晶高温合金中TCP相成核—生长机理的原位动态研究
-
批准号:52301178
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:夏万顺
-
依托单位:
NbZrTi基多主元合金中化学不均匀性对辐照行为的影响研究
-
批准号:12305290
-
项目类别:青年科学基金项目
-
资助金额:30.00万元
-
批准年份:2023
-
负责人:苏钲雄
-
依托单位:
眼表菌群影响糖尿病患者干眼发生的人群流行病学研究
-
批准号:82371110
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:邹海东
-
依托单位:
CuAgSe基热电材料的结构特性与构效关系研究
-
批准号:22375214
-
项目类别:面上项目
-
资助金额:50.00万元
-
批准年份:2023
-
负责人:周钲洋
-
依托单位:
镍基UNS N10003合金辐照位错环演化机制及其对力学性能的影响研究
-
批准号:12375280
-
项目类别:面上项目
-
资助金额:53.00万元
-
批准年份:2023
-
负责人:黄鹤飞
-
依托单位:
A study on prototype flexible multifunctional graphene foam-based sensing grid (柔性多功能石墨烯泡沫传感网格原型研究)
-
批准号:--
-
项目类别:--
-
资助金额:20万元
-
批准年份:2020
-
负责人:SAGAR RIZWAN UR REHMAN
-
依托单位:
基于大数据定量研究城市化对中国季节性流感传播的影响及其机理
-
批准号:82003509
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:雷浩
-
依托单位: