Deep architectures for statistical speech synthesis
Deep architectures for statistical speech synthesis
批准号:
EP/J002526/1
负责人:
Junichi Yamagishi
金额:
$94.44万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --
中文摘要
语音合成是将书面文本转换为语音输出。应用范围从电话对话系统到电脑游戏和临床应用。当前的语音合成系统可用的不同声音的范围非常有限。这是因为创建它们既复杂又昂贵。不幸的是,这对许多有趣的应用程序来说是一个大问题,包括我们在这项提案中关注的一个:为因运动神经元疾病和其他疾病而有发声问题的人提供辅助交流辅助工具。目前,这些人被迫使用声音不合适的设备,经常使用错误的口音,有时甚至是错误的性别!这对他们来说是一种阻碍,即使是与他们自己的家人交流,因为他们并不“拥有”声音,这也不能反映他们的身份。声音是身份不可或缺的一部分,我们正在创造一项技术,当人们的自然语言变得难以理解或根本无法说话时,他们可以用自己的声音进行交流。我们将开发的技术还具有许多其他应用:它将使语音合成器不仅能够调整说话人的身份,还可以调整许多其他属性。例如,调整说话努力将模拟人类谈话者在嘈杂环境中所做的事情,以使他们的语音更容易理解。我们的出发点是我们首创的一种技术,称为说话人自适应。事实证明,说话人自适应在基于少量记录语音的文本到语音合成系统的特征的灵活转换方面非常成功。它可以用于将语音的特征更改为不同的说话人或说话风格。然而,目前的方法没有使用任何关于语音的深入知识,也不能在类似的情况下进行推广。这比人类的言语产生要自然和灵活得多,在人类言语产生中,言语是由人类说话者简单地根据先前的经验来控制的。例如,与安静的环境相比,我们在嘈杂的环境中毫不费力地调整我们的语音,以提高可理解性。目前我们已经开创的适应技术是完全自动的,但它们不能以直接的方式将这种先验知识纳入其中。在一些初步工作中,我们开发了一个模型,其中包括关于语音发音的运动的信息:舌头、嘴唇等。然后,利用我们关于人类如何在噪声存在下改变他们的语音产生的知识,我们已经证明了在噪声中提高合成语音的可理解性是可能的。目前的建议是扩展和推广这项前期工作,以便将许多其他类型的关于人类语音的知识整合到这个模型中。我们将开发一种新的模型,允许我们包括更多关于语音是如何产生的信息,以及关于语音如何被感知以及外部因素(如背景噪音)如何影响语音的信息。这项技术的一个重要应用是为语音障碍患者(例如,运动神经元疾病引起的)创造个性化的语音合成。目前的发声技术对这些人不起作用,因为他们的语言通常已经混乱。我们的技术实际上可以纠正这一点,并产生听起来像人的语音,但比他们目前的自然语音更容易理解。我们已经制作了一个概念验证系统,证明了这一点是有效的。目前的提议将使广泛的人能够获得和负担得起这项技术。
英文摘要
Speech synthesis is the conversion of written text into speech output. Applications range from telephone dialogue systems to computer games and clinical applications. Current speech synthesis systems have a very limited range of difference voices available. This is because it is complex and expensive to create them.Unfortunately, that is a big problem for many interesting applications, including one we are focusing on in this proposal: assistive communication aids for people with vocal problems due to Motor Neurone Disease and other conditions. At the moment, these people are forced to use devices with inappropriate voices, very often in the wrong accent and sometimes even of the wrong sex! This is a disincentive for them to communicate, even with their own family, since they do not "own" the voice and it does not reflect their identity. The voice is an integral part of identity, and we are creating the technology to allow people to communicate in their own voice, when their natural speech has become hard to understand or they can no longer speak at all.The technology we will develop has a lot of other applications too: it will enable a speech synthesiser to adjust not only the speaker identity but many other properties too. For example, adjusting speaking effort will simulate what human talkers do in noisy conditions to make their speech more intelligible. Our starting point is a technique we have pioneered, called speaker adaptation.Speaker adaptation has proven to be highly successful in enabling the flexible transformation of the characteristics of a text-to-speech synthesis system, based on a small amount of recorded speech. It can be used for changing the characteristics of the speech to a different speaker or speaking style. However, current methods do not use any deep knowledge about speech and does not generalise across similar situations. This is considerably less natural and flexible than human speech production, in which speech is controlled by human talkers based simply on prior experience. For instance, we effortlessly adapt our speech in noisy environments, compared with quiet environments, in order to increase intelligibility. The current adaptation techniques that we have pioneered are completely automatic, but they do not enable this prior knowledge to be incorporated in a straightforward way.In some preliminary work, we have developed a model which includes information about the movement of the speech articulators: the tongue, lips and so on. Then, using our knowledge of how humans alter their speech production in the presence of noise (hyper- & hypo-articulation), we have demonstrated that it is possible to improve the intelligibility of synthetic speech in noise.The current proposal is to extend and generalise this preliminary work, in order to integrate many other types of knowledge about human speech into this model. We will develop a new model which allows us to include more information about how speech is produced, as well as information about how it is perceived and how external factors, such as background noise, affect speech.One important application of this technology is to create personalised speech synthesis for people with disordered speech (caused by Motor Neurone Disease, for example). Current technology for creating voices does not work for these people, because their speech is usually already disordered. Our technique can actually correct this, and produce speech which sounds like the person, but is more intelligible than their current natural speech. We have already produced a proof-of-concept system demonstrating that this works. The current proposal will make the technology available and affordable to a wide range of people.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Analysis of speaker clustering strategies for HMM-based speech synthesis
基于HMM的语音合成说话人聚类策略分析
DOI:
--
发表时间:
2012
期刊:
13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
影响因子:
--
作者:
[Dall R.]
通讯作者:
Dall R.
Glottal Spectral Separation for Speech Synthesis
用于语音合成的声门频谱分离
DOI:
10.1109/jstsp.2014.2307274
发表时间:
2014
期刊:
IEEE Journal of Selected Topics in Signal Processing
影响因子:
7.5
作者:
[Cabral J]
通讯作者:
Cabral J
DOI:
10.1109/icassp.2016.7472755
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou]
通讯作者:
Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou
A fixed dimension and perceptually based dynamic sinusoidal model of speech
固定维度和基于感知的动态正弦语音模型
DOI:
10.1109/icassp.2014.6854810
发表时间:
2014
期刊:
影响因子:
--
作者:
[Hu Q]
通讯作者:
Hu Q
DOI:
--
发表时间:
2012-12
期刊:
影响因子:
--
作者:
[MagdalenaAnnaKonkiewicz;Astrinak Maria;Yamagishi Junichi]
通讯作者:
MagdalenaAnnaKonkiewicz;Astrinak Maria;Yamagishi Junichi
共 9 条
海外基金