Using Local Phrase Dependency Structure Information in Neural Sequence-to-Sequence Speech Synthesis

Using Local Phrase Dependency Structure Information in Neural Sequence-to-Sequence Speech Synthesis
复制标题

在神经序列到序列语音合成中使用局部短语依赖结构信息

DOI:
10.1109/o-cocosda202152914.2021.9660456
复制
发表时间:
2021
期刊:
Proceeding of the Oriental COCOSDA 2021
影响因子:
--
通讯作者:
Satoshi Nakamura
Satoshi Nakamura
中科院分区:
--
文献类型:
--
作者:
Nobuyoshi Kaiki;Sakriani Sakti;Satoshi Nakamura

文献摘要

相似文献

我们介绍了端到端的文本到语音合成(TTS)与韵律符号,表示短语成分的基础上,本地句法依赖结构合成日语语音与自然韵律。我们提出了两个TTS模型:1)一个具有表示短语边界处的句法依赖性距离的韵律符号,以及2)另一个具有反映基于F0生成控制机制的短语和重音分量的叠加模型的韵律符号。使用这两个模型,我们观察到1)表示短语边界的停顿插入和2)F0在右分支边界处重置。为了验证这两个模型的有效性,对传统的模型只使用口音成分,我们进行了AB测试作为主观评价。我们的研究结果证实,与自然韵律,这反映了相应的意图的话语,合成语音生成使用本地短语依赖信息的句子和F0生成模型在日本的端到端的TTS。
We introduce end-to-end text-to-speech synthesis (TTS) with prosodic symbols that represent phrase components based on local syntactic dependency structures for synthesizing Japanese speech with natural prosody. We propose two TTS models: 1) one with prosodic symbols representing the syntactic dependency distance at the phrase boundaries and 2) another with prosodic symbols that reflect a superimposed model of the phrase and accent components based on an F0 generation control mechanism. Using these two models, we observed 1) pause insertion that indicates the phrase boundary and 2) F0 resetting at the right-branching boundaries. To verify the effectiveness of these two proposed models against the conventional model using only accent components, we conducted an AB test as a subjective evaluation. Our result confirmed that synthetic speech with natural prosody, which reflects the corresponding intention to the utterance, was generated using the local phrase dependency information of sentences and the F0 generation model in a Japanese end-to-end TTS.