Sequence-to-Sequence Learning for Deep Gaussian Process Based Speech Synthesis Using Self-Attention GP Layer

Sequence-to-Sequence Learning for Deep Gaussian Process Based Speech Synthesis Using Self-Attention GP Layer
复制标题

DOI:
10.21437/interspeech.2021-896
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Taiki Nakamura;Tomoki Koriyama;H. Saruwatari
Taiki Nakamura;Tomoki Koriyama;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Taiki Nakamura;Tomoki Koriyama;H. Saruwatari

文献摘要

相似文献

本文提出了一种基于深度高斯过程(DGP)和序列到序列(Seq2Seq)学习的语音合成方法,以实现高质量的端到端语音合成。由于贝叶斯学习和核回归,使用 DGP 的前馈和循环模型比深度神经网络 (DNN) 能产生更自然的合成语音。然而,此类 DGP 模型由独立模型、声学模型和持续时间模型的管道架构组成,并且需要高水平的文本处理专业知识。所提出的模型基于 Seq2Seq 学习,可以实现声学和持续时间模型的统一训练。编码器和解码器层由高斯过程回归(GPR)表示,参数被训练为贝叶斯模型。我们还提出了一种具有高斯过程的自注意力机制,可以有效地对编码器中的字符级输入进行建模。主观评价结果表明,所提出的 Seq2Seq-SA-DGP 可以比具有自注意力和循环结构的 DNN 合成更多的自然语音。此外,Seq2Seq-SA-DGP 减少了循环结构的平滑问题,并且在给定端到端系统的简单输入时是有效的。
This paper presents a speech synthesis method based on deep Gaussian process (DGP) and sequence-to-sequence (Seq2Seq) learning toward high-quality end-to-end speech synthesis. Feed-forward and recurrent models using DGP are known to produce more natural synthetic speech than deep neural networks (DNNs) because of Bayesian learning and kernel regression. However, such DGP models consist of a pipeline architecture of independent models, acoustic and duration models, and require a high level of expertise in text processing. The proposed model is based on Seq2Seq learning, which enables a unified training of acoustic and duration models. The encoder and decoder layers are represented by Gaussian process regressions (GPRs) and the parameters are trained as a Bayesian model. We also propose a self-attention mechanism with Gaussian processes to effectively model character-level input in the encoder. The subjective evaluation results show that the proposed Seq2Seq-SA-DGP can synthesize more natural speech than DNNs with self-attention and recurrent structures. Besides, Seq2Seq-SA-DGP reduces the smoothing problems of recurrent structures and is effective when a simple input for an end-to-end system is given.