Direct Modeling of Frequency Spectra and Waveform Generation Based on Phase Recovery for DNN-Based Speech Synthesis

Direct Modeling of Frequency Spectra and Waveform Generation Based on Phase Recovery for DNN-Based Speech Synthesis
复制标题

DOI:
10.21437/interspeech.2017-488
复制
发表时间:
2017-08
期刊:
--
影响因子:
--
通讯作者:
Shinji Takaki;H. Kameoka;J. Yamagishi
Shinji Takaki;H. Kameoka;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Shinji Takaki;H. Kameoka;J. Yamagishi

文献摘要

相似文献

在使用高质量声码器的统计参数语音合成(SPSS)系统中,从语言特征中预测诸如梅尔倒谱系数、fi指数和F0的声学特征,以便利用声码器来生成语音波形。然而,所生成的语音波形通常不会受到质量恶化的影响,例如由于使用声码器而引起的嗡嗡声。虽然已经研究了几种尝试,如改进激励模型来缓解这一问题,但如果spss系统是基于声码器的,则很难完全避免这一问题。为了克服这个问题,最近有人尝试直接对波形样本进行建模。卓越的性能已经得到证明,但计算时间和延迟仍然是问题。为了构建另一种既不需要声码器又不需要计算爆炸的基于DNN的语音合成器,我们研究了基于相位恢复的频谱直接建模和波形产生方法。在该框架下,通过基于离散神经网络的声学模型直接预测包含来自F0的谐波信息的短时傅立叶变换频谱幅度,并使用GRIFfin和LIM的方法来恢复相位和产生波形。实验结果表明,该系统合成的语音没有杂音,性能优于传统的声码器语音合成系统。
In statistical parametric speech synthesis (SPSS) systems using the high-quality vocoder, acoustic features such as mel-cepstrum coefficients and F 0 are predicted from linguistic features in order to utilize the vocoder to generate speech waveforms. However, the generated speech waveform generally suf-fers from quality deterioration such as buzziness caused by uti-lizing the vocoder. Although several attempts such as improv-ing an excitation model have been investigated to alleviate the problem, it is difficult to completely avoid it if the SPSS system is based on the vocoder. To overcome this problem, there have recently been attempts to directly model waveform samples. Superior performance has been demonstrated, but com-putation time and latency are still issues. With the aim to construct another type of DNN-based speech synthesizer with nei-ther the vocoder nor computational explosion, we investigated direct modeling of frequency spectra and waveform generation based on phase recovery. In this framework, STFT spectral amplitudes that include harmonics information derived from F 0 are directly predicted through a DNN-based acoustic model and we use Griffin and Lim’s approach to recover phase and generate waveforms. The experimental results showed that the proposed system synthesized speech without buzziness and outperformed one generated from a conventional system using the vocoder.