Deep architectures for statistical speech synthesis
Deep architectures for statistical speech synthesis
批准号:
EP/J002526/1
负责人:
Junichi Yamagishi
金额:
$94.44万
依托单位:
依托单位国家:
英国
项目类别:
Fellowship
财政年份:
2011
资助国家:
英国
项目状态:
已结题
起止时间:
2011 至 --
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Speech synthesis is the conversion of written text into speech output. Applications range from telephone dialogue systems to computer games and clinical applications. Current speech synthesis systems have a very limited range of difference voices available. This is because it is complex and expensive to create them.Unfortunately, that is a big problem for many interesting applications, including one we are focusing on in this proposal: assistive communication aids for people with vocal problems due to Motor Neurone Disease and other conditions. At the moment, these people are forced to use devices with inappropriate voices, very often in the wrong accent and sometimes even of the wrong sex! This is a disincentive for them to communicate, even with their own family, since they do not "own" the voice and it does not reflect their identity. The voice is an integral part of identity, and we are creating the technology to allow people to communicate in their own voice, when their natural speech has become hard to understand or they can no longer speak at all.The technology we will develop has a lot of other applications too: it will enable a speech synthesiser to adjust not only the speaker identity but many other properties too. For example, adjusting speaking effort will simulate what human talkers do in noisy conditions to make their speech more intelligible. Our starting point is a technique we have pioneered, called speaker adaptation.Speaker adaptation has proven to be highly successful in enabling the flexible transformation of the characteristics of a text-to-speech synthesis system, based on a small amount of recorded speech. It can be used for changing the characteristics of the speech to a different speaker or speaking style. However, current methods do not use any deep knowledge about speech and does not generalise across similar situations. This is considerably less natural and flexible than human speech production, in which speech is controlled by human talkers based simply on prior experience. For instance, we effortlessly adapt our speech in noisy environments, compared with quiet environments, in order to increase intelligibility. The current adaptation techniques that we have pioneered are completely automatic, but they do not enable this prior knowledge to be incorporated in a straightforward way.In some preliminary work, we have developed a model which includes information about the movement of the speech articulators: the tongue, lips and so on. Then, using our knowledge of how humans alter their speech production in the presence of noise (hyper- & hypo-articulation), we have demonstrated that it is possible to improve the intelligibility of synthetic speech in noise.The current proposal is to extend and generalise this preliminary work, in order to integrate many other types of knowledge about human speech into this model. We will develop a new model which allows us to include more information about how speech is produced, as well as information about how it is perceived and how external factors, such as background noise, affect speech.One important application of this technology is to create personalised speech synthesis for people with disordered speech (caused by Motor Neurone Disease, for example). Current technology for creating voices does not work for these people, because their speech is usually already disordered. Our technique can actually correct this, and produce speech which sounds like the person, but is more intelligible than their current natural speech. We have already produced a proof-of-concept system demonstrating that this works. The current proposal will make the technology available and affordable to a wide range of people.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Analysis of speaker clustering strategies for HMM-based speech synthesis
基于HMM的语音合成说话人聚类策略分析
DOI:
--
发表时间:
2012
期刊:
13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
影响因子:
--
作者:
[Dall R.]
通讯作者:
Dall R.
Glottal Spectral Separation for Speech Synthesis
用于语音合成的声门频谱分离
DOI:
10.1109/jstsp.2014.2307274
发表时间:
2014
期刊:
IEEE Journal of Selected Topics in Signal Processing
影响因子:
7.5
作者:
[Cabral J]
通讯作者:
Cabral J
DOI:
10.1109/icassp.2016.7472755
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou]
通讯作者:
Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou
A fixed dimension and perceptually based dynamic sinusoidal model of speech
固定维度和基于感知的动态正弦语音模型
DOI:
10.1109/icassp.2014.6854810
发表时间:
2014
期刊:
影响因子:
--
作者:
[Hu Q]
通讯作者:
Hu Q
DOI:
--
发表时间:
2012-12
期刊:
影响因子:
--
作者:
[MagdalenaAnnaKonkiewicz;Astrinak Maria;Yamagishi Junichi]
通讯作者:
MagdalenaAnnaKonkiewicz;Astrinak Maria;Yamagishi Junichi
共 9 条
海外基金