On Appropriateness and Estimation of the Emotion of Synthesized Response Speech in a Spoken Dialogue System

On Appropriateness and Estimation of the Emotion of Synthesized Response Speech in a Spoken Dialogue System
复制标题

DOI:
10.1007/978-3-319-21380-4_126
复制
发表时间:
2015-08
期刊:
--
影响因子:
--
通讯作者:
Taketo Kase;Takashi Nose;Akinori Ito
Taketo Kase;Takashi Nose;Akinori Ito
中科院分区:
其他
文献类型:
--
作者:
Taketo Kase;Takashi Nose;Akinori Ito

文献摘要

相似文献

在口语对话系统中,话语的情感等副语言特征对于生成更好的响应话语与其语言内容一样重要。在本研究中,我们进行了一个实验,以揭示在对话系统中的情感语音合成的效果,并调查了什么方法是有效的,给情感的合成语音。首先,我们进行了一个实验,其中一个代理与各种情绪的语音与用户交谈,和适当的情绪进行了评估。正如预期的那样,当我们适当地添加情感时,用户对代理的印象更好。接下来,我们研究了自动估计系统响应的情感的方法,我们发现最好的方法是给出与用户先前话语相同的情感,而不管系统话语的内容。
Paralinguistic features such as emotion of an utterance is as important as its linguistic content for generating better response utterances in spoken dialog systems. In this research, we carried out an experiment to reveal the effect of emotional speech synthesis in a dialogue system, and investigated what method was effective for giving emotion to the synthetic speech. Firstly, we carried out an experiment where an agent with various emotional speech talked to the user, and the appropriateness of the emotion was evaluated. As expected, users had better impression on the agent when we added emotion appropriately. Next, we examined methods of automatic estimation of emotion for the system’s response, and we found that the best method was to give the same emotion as the user’s previous utterance regardless of the content of the system’s utterance.