Speaker-independent style conversion for HMM-based expressive speech synthesis

Speaker-independent style conversion for HMM-based expressive speech synthesis
复制标题

DOI:
10.1109/icassp.2013.6639195
复制
发表时间:
2013-05
期刊:
2013 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Hiroki Kanagawa;Takashi Nose;Takao Kobayashi
Hiroki Kanagawa;Takashi Nose;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
Hiroki Kanagawa;Takashi Nose;Takao Kobayashi

文献摘要

相似文献

在基于隐马尔可夫模型的语音合成中,提出了一种由目标说话人的中性风格语音生成目标说话人表达风格模型的方法。该技术基于使用线性变换的风格适配,其中使用多个说话人的中性和目标风格的语音数据对预先估计与说话人无关的变换矩阵。通过将得到的变换矩阵应用于新的说话人的中性风格模型,我们可以将声学模型的风格表现力转换为目标风格,而不需要准备任何目标风格的语音。此外,我们还将说话人自适应训练(SAT)框架引入到变换估计中,以减小说话人之间的声学差异。我们从自然度、说话人相似度和风格重现性三个方面对风格转换的性能进行了主观评价。
This paper proposes a technique for creating target speaker's expressive-style model from the target speaker's neutral style speech in HMM-based speech synthesis. The technique is based on the style adaptation using linear transforms where speaker-independent transformation matrices are estimated in advance using pairs of neutral and target-style speech data of multiple speakers. By applying the obtained transformation matrices to a new speaker's neutral-style model, we can convert the style expressivity of the acoustic model to the target style without preparing any target-style speech of the speaker. In addition, we introduce a speaker adaptive training (SAT) framework into the transform estimation to reduce the acoustic difference among speakers. We subjectively evaluate the performance of the style conversion in terms of the naturalness, speaker similarity, and style reproducibility.