An RNN-Based Quantized F0 Model with Multi-Tier Feedback Links for Text-to-Speech Synthesis

An RNN-Based Quantized F0 Model with Multi-Tier Feedback Links for Text-to-Speech Synthesis
复制标题

DOI:
10.21437/interspeech.2017-246
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Xin Wang;Shinji Takaki;J. Yamagishi
Xin Wang;Shinji Takaki;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Xin Wang;Shinji Takaki;J. Yamagishi

文献摘要

相似文献

提出了一种基于递归神经网络的语音合成F0模型,该模型能在给定文本特征的情况下生成F0轮廓。与相关的F0模型相比,所提出的模型旨在学习多个级别的F0轮廓的时间相关性。帧级相关性通过反馈前一帧的F0输出作为当前帧的附加输入来覆盖;同时,类似地建模长时间跨度上的相关性,但是通过使用在音素和音节上聚合的F0特征。另一个区别是,所提出的模型的输出不是插值的连续值F0轮廓,而是一个序列的离散符号,包括量化的F0电平和一个符号的清音条件。通过使用离散的F0符号,该模型避免了阿尔蒂人工插值F0曲线的影响。实验表明,使用dropout策略训练的F0模型生成的平滑F0轮廓比基线RNN模型具有相对更好的感知质量。
A recurrent-neural-network-based F0 model for text-to-speech (TTS) synthesis that generates F0 contours given textual features is proposed. In contrast to related F0 models, the proposed one is designed to learn the temporal correlation of F0 contours at multiple levels. The frame-level correlation is covered by feeding back the F0 output of the previous frame as the additional input of the current frame; meanwhile, the correlation over long-time spans is similarly modeled but by using F0 features aggregated over the phoneme and syllable. Another difference is that the output of the proposed model is not the interpolated continuous-valued F0 contour but rather a sequence of discrete symbols, including quantized F0 levels and a symbol for the unvoiced condition. By using the discrete F0 symbols, the proposed model avoids the influence of artificially interpolated F0 curves. Experiments demonstrated that the proposed F0 model, which was trained using a dropout strategy, generated smooth F0 contours with relatively better perceived quality than those from baseline RNN models.