Vocal attractiveness of statistical speech synthesisers

Vocal attractiveness of statistical speech synthesisers
复制标题

统计语音合成器的声音吸引力

DOI:
--
复制
发表时间:
2011
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
Simon King
Simon King
中科院分区:
--
文献类型:
--
作者:
Sandra Andraszewicz;J. Yamagishi;Simon King

文献摘要

被引文献

相似文献

我们之前对基于说话人自适应HMM的语音合成方法的分析表明,平均语音可以获得比任何单个自适应语音更高的主观分数的原因有两个:1)模型自适应会使语音质量与变换“移动”的距离成比例地降低,以及2)与语音吸引力相关的心理声学效应。本文是该分析的后续,旨在将这些影响分离出来。我们最新的感知实验专注于吸引力,使用平均语音和扬声器依赖的语音,而无需模型转换,并表明使用多个扬声器创建语音可以提高平滑度(通过谐波噪声比测量),减少与最终语音的log F0-F1空间中的平均语音的距离,从而使其在分段水平上更具吸引力。然而,这在超音段或句子层面上被削弱或被推翻。
Our previous analysis of speaker-adaptive HMM-based speech synthesis methods suggested that there are two possible reasons why average voices can obtain higher subjective scores than any individual adapted voice: 1) model adaptation degrades speech quality proportionally to the distance ‘moved’ by the transforms, and 2) psychoacoustic effects relating to the attractiveness of the voice. This paper is a follow-on from that analysis and aims to separate these effects out. Our latest perceptual experiments focus on attractiveness, using average voices and speaker-dependent voices without model transformation, and show that using several speakers to create a voice improves smoothness (measured by Harmonics-to-Noise Ratio), reduces distance from the the average voice in the log F0-F1 space of the final voice and hence makes it more attractive at the segmental level. However, this is weakened or overridden at supra-segmental or sentence levels.