Conversational End-to-End TTS for Voice Agents

Conversational End-to-End TTS for Voice Agents
复制标题

语音代理的会话式端到端 TTS

DOI:
10.1109/slt48900.2021.9383460
复制
发表时间:
2020
期刊:
2021 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Lei Xie
Lei Xie
中科院分区:
--
文献类型:
--
作者:
Haohan Guo;Shaofei Zhang;F. Soong;Lei He;Lei Xie

文献摘要

参考文献

被引文献

相似文献

端到端的神经TTS在阅读风格的语音合成上取得了优异的性能。然而,由于语料库和建模能力的限制,构建高质量的会话式文语转换系统仍然是一个挑战。本研究的目的是在序列到序列建模框架下为语音代理构建一个会话式TTS。我们首先构建了一个为语音代理设计的自发会话语音语料库,并采用了一种新的录音方案,既保证了录音质量,又保证了会话的风格。其次,我们提出了一种会话上下文感知的端到端TTS方法,该方法采用辅助编码器和会话上下文编码器来专门加强有关会话中当前话语及其上下文的信息。实验结果表明,该方法产生更自然的韵律根据会话的背景,具有显着的偏好增益在话语级和会话级。此外,我们还发现该模型能够表达一些自发的行为,如填充和重复的话,这使得会话风格更真实。
End-to-end neural TTS has achieved excellent performance on reading style speech synthesis. However, it is still a challenge to build a high-quality conversational TTS due to the limitations of corpus and modeling capability. This study aims at building a conversational TTS for a voice agent under sequence to sequence modeling framework. We firstly construct a spontaneous conversational speech corpus well designed for the voice agent with a new recording scheme ensuring both recording quality and conversational speaking style. Secondly, we propose a conversation context-aware end-to-end TTS approach that employs an auxiliary encoder and a conversational context encoder to specifically reinforce the information about the current utterance and its context in a conversation as well. Experimental results show that the proposed approach produces more natural prosody in accordance with the conversational context, with significant preference gains at both utterance-level and conversation-level. Moreover, we find that the model has the ability to express some spontaneous behaviors like fillers and repeated words, which makes the conversational speaking style more realistic.
DOI: --
发表时间: --
期刊:
影响因子: --
作者:
Y. Sagisaka;Y. Greenberg;K. Li;M. Zhu;M. Tsuzaki;H. Kato
通讯作者: H. Kato
扩展上下文在基于 HMM 的自发会话语音合成中的应用
DOI: --
发表时间: 2011
期刊: Proceedings of the 12th Annual Conference of the International Speech Communication Association, INTERSPEECH 2011
影响因子: --
作者:
Ken Ichikawa;Satoru Tsuge;Norihide Kitaoka;Kazuya Takeda;Kenji Kita;Tomoki Koriyama
通讯作者: Tomoki Koriyama