A classifier-based target cost for unit selection speech synthesis trained on perceptual data

A classifier-based target cost for unit selection speech synthesis trained on perceptual data
复制标题

DOI:
10.21437/interspeech.2010-72
复制
发表时间:
2010
期刊:
--
影响因子:
--
通讯作者:
V. Strom;Simon King
V. Strom;Simon King
中科院分区:
其他
文献类型:
--
作者:
V. Strom;Simon King

文献摘要

被引文献

相似文献

我们的目标是自动学习一个感知最优的目标成本函数的单元选择语音合成器。我们在这里采取的方法是训练一个分类器的人类感知判断的合成语音。分类器的输出用于进行简单的三向区分,而不是估计连续值的成本。为了收集必要的感知数据,我们合成了145,137个短句,关闭了通常的目标成本,因此搜索仅由连接成本驱动。然后,我们选择了7200个连接最好的句子,并要求60名听众对它们进行评判,提供他们对每个音节的评分。由此,我们得出了每一款智能手机的评级。使用与我们传统的目标成本函数相同的上下文特征作为输入,我们在这些人类感知评级上训练了一个分类器。我们合成了两组测试句子与我们的标准目标成本和新的目标成本的基础上的分类。A/B偏好测试表明,基于分类器的目标成本,这是完全自动地从适量的感知数据学习,几乎是一样好,我们仔细和专业调整的标准目标成本。
Our goal is to automatically learn a perceptually-optimal target cost function for a unit selection speech synthesiser. The approach we take here is to train a classifier on human perceptual judgements of synthetic speech. The output of the classifier is used to make a simple three-way distinction rather than to estimate a continuously-valued cost. In order to collect the necessary perceptual data, we synthesised 145,137 short sentences with the usual target cost switched off, so that the search was driven by the join cost only. We then selected the 7200 sentences with the best joins and asked 60 listeners to judge them, providing their ratings for each syllable. From this, we derived a rating for each demiphone. Using as input the same context features employed in our conventional target cost function, we trained a classifier on these human perceptual ratings. We synthesised two sets of test sentences with both our standard target cost and the new target cost based on the classifier. A/B preference tests showed that the classifier-based target cost, which was learned completely automatically from modest amounts of perceptual data, is almost as good as our carefullyand expertly-tuned standard target cost.