Naturalness and rapport in a pitch adaptive learning companion

Naturalness and rapport in a pitch adaptive learning companion
复制标题

音调自适应学习伙伴的自然性和融洽感

DOI:
10.1109/asru.2015.7404781
复制
发表时间:
2015
期刊:
2015 IEEE Workshop on Automatic Speech Recognition and Understanding (ASRU)
影响因子:
--
通讯作者:
Erin Walker
Erin Walker
中科院分区:
--
文献类型:
--
作者:
Nichola Lubold;Heather Pon;Erin Walker

文献摘要

被引文献

相似文献

在人与人的互动中经常观察到,夹带是一种社会现象,说话者在谈话过程中变得更加相似。当个体调整其声学韵律语音特征(例如音调和强度)时,就会发生声学韵律夹带。与沟通的成功、自然性和会话流畅以及融洽等社会变量相关,自动引导的对话系统有可能通过增加融洽、自然和会话流畅来改善言语互动。在像学习伙伴这样的应用程序中,这种社交响应对话系统可以改善学习和动机。然而,目前尚不清楚如何在自动对话系统中产生夹带,从而产生人与人对话中看到的效果。在本文中,我们迈出了实现可引导的语音对话系统的第一步。我们基于对人类夹带的分析提出了三种音调适应方法,并设计和实现了一个可以自适应地操纵文本到语音输出的音调的系统。我们发现融洽感与不同形式的音调适应之间存在明显的关系。某些适应被认为明显更加自然和融洽。最终,通过将文本到语音输出的音调轮廓移动用户的平均音调来进行调整,从而获得最高报告的融洽度和自然度测量值。
Observed frequently in human-human interactions, entrainment is a social phenomenon in which speakers become more like each other over the course of a conversation. Acoustic-prosodic entrainment occurs when individuals adapt their acoustic-prosodic speech features, such as pitch and intensity. Correlated with communicative success, naturalness, and conversational flow as well as social variables such as rapport, a dialogue system which automatically entrains has the potential to improve verbal interactions by increasing rapport, naturalness, and conversational flow. In an application like the learning companion, such a socially responsive dialogue system may improve learning and motivation. However, it is not clear how to produce entrainment in an automatic dialogue system in ways that produce the effects seen in human-human dialogue. In this paper, we take the first steps towards implementing a spoken dialogue system which can entrain. We propose three methods of pitch adaptation based on analysis of human entrainment, and design and implement a system which can manipulate the pitch of text-to-speech output adaptively. We find a clear relationship between perceptions of rapport and different forms of pitch adaptations. Certain adaptations are perceived as significantly more natural and rapport-like. Ultimately, adapting by shifting the pitch contour of the text-to-speech output by the mean pitch of the user results in the highest reported measures of rapport and naturalness.