Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History

Acoustic Modeling for End-to-End Empathetic Dialogue Speech Synthesis Using Linguistic and Prosodic Contexts of Dialogue History
复制标题

DOI:
10.48550/arxiv.2206.08039
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Yuto Nishimura;Yuki Saito;Shinnosuke Takamichi;Kentaro Tachibana;H. Saruwatari
Yuto Nishimura;Yuki Saito;Shinnosuke Takamichi;Kentaro Tachibana;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
Yuto Nishimura;Yuki Saito;Shinnosuke Takamichi;Kentaro Tachibana;H. Saruwatari

文献摘要

相似文献

我们提出了一个端到端的共情对话语音合成(DSS)模型,该模型考虑了对话历史的语言和韵律背景。共情是人类在对话中进入对话者内心的主动尝试,而共情决策支持系统是在口语对话系统中实现这一行为的一种技术。我们的模型是根据语言和韵律的历史特征来预测适当的对话语境的。因此,它可以看作是传统的基于语言特征的对话历史建模的扩展。为了有效地训练共情决策支持系统模型,我们研究了1)使用大型语音语料库预训练的自监督学习模型,2)使用对话上下文嵌入预测当前话语的韵律嵌入的风格引导训练,3)结合文本和语音模态的跨模态关注,以及4)基于句子的嵌入来实现细粒度韵律建模,而不是基于话语的建模。评价结果表明:1)单纯考虑对话历史的韵律语境并不能提高共情决策支持系统的语音质量;2)引入风格引导训练和基于句子的嵌入建模比传统方法获得更高的语音质量。
We propose an end-to-end empathetic dialogue speech synthesis (DSS) model that considers both the linguistic and prosodic contexts of dialogue history. Empathy is the active attempt by humans to get inside the interlocutor in dialogue, and empathetic DSS is a technology to implement this act in spoken dialogue systems. Our model is conditioned by the history of linguistic and prosody features for predicting appropriate dialogue context. As such, it can be regarded as an extension of the conventional linguistic-feature-based dialogue history modeling. To train the empathetic DSS model effectively, we investigate 1) a self-supervised learning model pretrained with large speech corpora, 2) a style-guided training using a prosody embedding of the current utterance to be predicted by the dialogue context embedding, 3) a cross-modal attention to combine text and speech modalities, and 4) a sentence-wise embedding to achieve fine-grained prosody modeling rather than utterance-wise modeling. The evaluation results demonstrate that 1) simply considering prosodic contexts of the dialogue history does not improve the quality of speech in empathetic DSS and 2) introducing style-guided training and sentence-wise embedding modeling achieves higher speech quality than that by the conventional method.