Direct Modeling of Frequency Spectra and Waveform Generation Based on Phase Recovery for DNN-Based Speech Synthesis
Direct Modeling of Frequency Spectra and Waveform Generation Based on Phase Recovery for DNN-Based Speech Synthesis
复制标题
DOI:
10.21437/interspeech.2017-488
复制
发表时间:
2017-08
期刊:
影响因子:
--
通讯作者:
Shinji Takaki;H. Kameoka;J. Yamagishi
中科院分区:
文献类型:
--
作者:
Shinji Takaki;H. Kameoka;J. Yamagishi
In statistical parametric speech synthesis (SPSS) systems using the high-quality vocoder, acoustic features such as mel-cepstrum coefficients and F 0 are predicted from linguistic features in order to utilize the vocoder to generate speech waveforms. However, the generated speech waveform generally suf-fers from quality deterioration such as buzziness caused by uti-lizing the vocoder. Although several attempts such as improv-ing an excitation model have been investigated to alleviate the problem, it is difficult to completely avoid it if the SPSS system is based on the vocoder. To overcome this problem, there have recently been attempts to directly model waveform samples. Superior performance has been demonstrated, but com-putation time and latency are still issues. With the aim to construct another type of DNN-based speech synthesizer with nei-ther the vocoder nor computational explosion, we investigated direct modeling of frequency spectra and waveform generation based on phase recovery. In this framework, STFT spectral amplitudes that include harmonics information derived from F 0 are directly predicted through a DNN-based acoustic model and we use Griffin and Lim’s approach to recover phase and generate waveforms. The experimental results showed that the proposed system synthesized speech without buzziness and outperformed one generated from a conventional system using the vocoder.