Autoregressive Neural F0 Model for Statistical Parametric Speech Synthesis

Autoregressive Neural F0 Model for Statistical Parametric Speech Synthesis
复制标题

DOI:
10.1109/taslp.2018.2828650
复制
发表时间:
2018-08
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Xin Wang;Shinji Takaki;J. Yamagishi
Xin Wang;Shinji Takaki;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Xin Wang;Shinji Takaki;J. Yamagishi

文献摘要

相似文献

递归神经网络(RNN)已被成功地用作文本到语音合成的基频(F0)模型。然而,本文表明,正常的RNN可能没有考虑帧间F0数据的统计相关性,因此仅当从模型中采样F0值时才产生噪声F0轮廓。更好的模型可以考虑当前F0数据对先前帧的F0数据的因果依赖性。一个这样的模型是我们最近提出的浅自回归(AR)递归混合密度网络(SAR)。然而,正如这项研究所表明的那样,合成孔径雷达相当于可训练的线性滤波器和传统的神经网络的组合。对于F0建模来说,它仍然是薄弱的。为了更好地模拟F0轮廓的时间相关性,我们提出了一种深度AR模型(DAR)。在RNN的基础上,该DAR通过RNN传播前一帧的F0值,这允许实现非线性AR相关性。我们还提出了用于DAR的F0量化和数据丢弃策略。在日语语料上的实验表明,这种基于随机抽样的DAR方法可以生成合适的F0轮廓,这是基线RNN和SAR所不能做到的。当DAR和其他实验模型使用传统的基于均值的生成方法时,DAR生成了准确且不太平滑的F0轮廓,并在主观评估测试中获得了更好的平均意见分数。
Recurrent neural networks (RNNs) have been successfully used as fundamental frequency (F0) models for text-to-speech synthesis. However, this paper showed that a normal RNN may not take into account the statistical dependency of the F0 data across frames and consequently only generate noisy F0 contours when F0 values are sampled from the model. A better model may take into account the causal dependency of the current F0 datum on the previous frames’ F0 data. One such model is the shallow autoregressive (AR) recurrent mixture density network (SAR) that we recently proposed. However, as this study showed, an SAR is equivalent to the combination of trainable linear filters and a conventional RNN. It is still weak for F0 modeling. To better model the temporal dependency in F0 contours, we propose a deep AR model (DAR). On the basis of an RNN, this DAR propagates the previous frame's F0 value through the RNN, which allows nonlinear AR dependency to be achieved. We also propose F0 quantization and data dropout strategies for the DAR. Experiments on a Japanese corpus demonstrated that this DAR can generate appropriate F0 contours by using the random-sampling-based generation method, which is impossible for the baseline RNN and SAR. When a conventional mean-based generation method was used in the proposed DAR and other experimental models, the DAR generated accurate and less oversmoothed F0 contours and achieved a better mean-opinion-score in a subjective evaluation test.