Using Local Phrase Dependency Structure Information in Neural Sequence-to-Sequence Speech Synthesis
Using Local Phrase Dependency Structure Information in Neural Sequence-to-Sequence Speech Synthesis
复制标题
在神经序列到序列语音合成中使用局部短语依赖结构信息
DOI:
10.1109/o-cocosda202152914.2021.9660456
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Satoshi Nakamura
中科院分区:
文献类型:
--
作者:
Nobuyoshi Kaiki;Sakriani Sakti;Satoshi Nakamura
We introduce end-to-end text-to-speech synthesis (TTS) with prosodic symbols that represent phrase components based on local syntactic dependency structures for synthesizing Japanese speech with natural prosody. We propose two TTS models: 1) one with prosodic symbols representing the syntactic dependency distance at the phrase boundaries and 2) another with prosodic symbols that reflect a superimposed model of the phrase and accent components based on an F0 generation control mechanism. Using these two models, we observed 1) pause insertion that indicates the phrase boundary and 2) F0 resetting at the right-branching boundaries. To verify the effectiveness of these two proposed models against the conventional model using only accent components, we conducted an AB test as a subjective evaluation. Our result confirmed that synthetic speech with natural prosody, which reflects the corresponding intention to the utterance, was generated using the local phrase dependency information of sentences and the F0 generation model in a Japanese end-to-end TTS.