A Hierarchical Encoder-Decoder Model for Statistical Parametric Speech Synthesis

A Hierarchical Encoder-Decoder Model for Statistical Parametric Speech Synthesis
复制标题

DOI:
10.21437/interspeech.2017-628
复制
发表时间:
2017-05
期刊:
--
影响因子:
--
通讯作者:
S. Ronanki;O. Watts;Simon King
S. Ronanki;O. Watts;Simon King
中科院分区:
其他
文献类型:
--
作者:
S. Ronanki;O. Watts;Simon King

文献摘要

被引文献

相似文献

目前使用神经网络进行统计参数语音合成的方法通常要求输入具有与输出相同的时间分辨率,通常每5 ms一帧,或者在某些情况下以波形采样率。因此,有必要在输入处构造高度冗余的帧级(或样本级)语言特征。本文提出了使用一个分层的编码器-解码器模型来执行序列到序列的回归的方式,以输入的语言特征在其原始的时间尺度,并保留单词,音节和电话之间的关系。所提出的模型旨在比传统架构更有效地利用超节段特征,并且具有计算效率。实验进行了韵律变化的有声读物材料,因为使用超音段的功能被认为是特别重要的,在这种情况下。客观测量和主观听力测试的结果(要求听众专注于韵律)表明,所提出的方法的性能明显优于要求语言输入处于声学帧速率的传统架构。我们提供了代码和配方,使我们的系统能够使用Merlin工具包进行复制。
Current approaches to statistical parametric speech synthesis using Neural Networks generally require input at the same temporal resolution as the output, typically a frame every 5ms, or in some cases at waveform sampling rate. It is therefore necessary to fabricate highly-redundant frame-level (or sample-level) linguistic features at the input. This paper proposes the use of a hierarchical encoder-decoder model to perform the sequence-to-sequence regression in a way that takes the input linguistic features at their original timescales, and preserves the relationships between words, syllables and phones. The proposed model is designed to make more effective use of supra-segmental features than conventional architectures, as well as being computationally efficient. Experiments were conducted on prosodically-varied audiobook material because the use of supra-segmental features is thought to be particularly important in this case. Both objective measures and results from subjective listening tests, which asked listeners to focus on prosody, show that the proposed method performs significantly better than a conventional architecture that requires the linguistic input to be at the acoustic frame rate. We provide code and a recipe to enable our system to be reproduced using the Merlin toolkit.