Neural Spoken-Response Generation Using Prosodic and Linguistic Context for Conversational Systems

Neural Spoken-Response Generation Using Prosodic and Linguistic Context for Conversational Systems
复制标题

DOI:
10.21437/interspeech.2021-381
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Yoshihiro Yamazaki;Yuya Chiba;Takashi Nose;Akinori Ito
Yoshihiro Yamazaki;Yuya Chiba;Takashi Nose;Akinori Ito
中科院分区:
其他
文献类型:
--
作者:
Yoshihiro Yamazaki;Yuya Chiba;Takashi Nose;Akinori Ito

文献摘要

被引文献

相似文献

口语对话系统已经广泛应用于日常生活中。这样的系统必须与用户进行社交互动,才能真正成为人类的合作伙伴。在最近的对话系统的研究中,神经反应生成导致自然反应生成。然而,这些研究没有考虑会话现象的声学方面,如韵律的适应。我们提出了一个口语反应生成模型,扩展了神经会话模型来处理音高控制信号。我们提出的模型使用人类之间的多模态对话进行训练。所产生的音调控制信号被输入到语音合成系统以控制合成语音的音调。我们的实验表明,所提出的系统可以生成合成语音与适当的F0轮廓作为话语的上下文中相比,没有音高控制的系统的输出,虽然语言生成仍然是一个问题。
Spoken dialogue systems have become widely used in daily life. Such a system must interact with the user socially to truly operate as a partner with humans. In studies of recent dialogue systems, neural response generation led to natural response generation. However, these studies have not considered the acoustic aspects of conversational phenomena, such as the adaptation of prosody. We propose a spoken-response generation model that extends a neural conversational model to deal with pitch con-trol signals. Our proposed model is trained using multimodal dialogue between humans. The generated pitch control signals are input to a speech synthesis system to control the pitch of synthesized speech. Our experiment shows that the proposed system can generate synthesized speech with an appropriate F0 contour as an utterance in context compared to the output of a system without pitch control, although language generation remains an issue.