Prosodic alignment toward emotionally expressive speech: Comparing human and Alexa model talkers

Prosodic alignment toward emotionally expressive speech: Comparing human and Alexa model talkers
复制标题

DOI:
10.1016/j.specom.2021.10.003
复制
发表时间:
2021-10-30
影响因子:
3.2
通讯作者:
Zellou, Georgia
Zellou, Georgia
中科院分区:
计算机科学3区
文献类型:
--
作者:
Cohn, Michelle;Predeck, Kristin;Zellou, Georgia

文献摘要

被引文献

相似文献

这项研究测试了两种类型的对话者(人类和语音激活的人工智能(语音AI)助手)产生的情感表达韵律是否与个人的声音一致。参与者完成了一个感叹词的单词阴影实验(例如,“Awesome”)由人类声音和语音AI系统(亚马逊的Alexa)生成的声音以情感中立和富有表现力的韵律产生。结果显示,增加参与者的词的持续时间,平均f0,和f0变化的情绪表达,一致的增加对齐一般的“积极情绪”的讲话风格。说话者类别(人类与语音-AI)的情感对齐的微小差异与模型说话者的声音差异平行,这表明参与者反映了他们听到的声学效果。人类和语音AI说话者对情感的相似反应支持了无中介情感对齐的解释,以及计算机拟人化:人们对这两种类型的对话者都施加了情感中介行为。虽然参与者性别之间的差异很小,但女性和男性的总体模式相似,支持情感声音对齐的细微差别。
This study tests whether individuals vocally align toward emotionally expressive prosody produced by two types of interlocutors: a human and a voice-activated artificially intelligent (voice-AI) assistant. Participants completed a word shadowing experiment of interjections (e.g., "Awesome") produced in emotionally neutral and expressive prosodies by both a human voice and a voice generated by a voice-AI system (Amazon's Alexa). Results show increases in participants' word duration, mean f0, and f0 variation in response to emotional expressiveness, consistent with increased alignment toward a general 'positive-emotional' speech style. Small differences in emotional alignment by talker category (human vs. voice-AI) parallel the acoustic differences in the model talkers' productions, suggesting that participants are mirroring the acoustics they hear. The similar responses to emotion in both a human and voice-AI talker support accounts of unmediated emotional alignment, as well as computer personification: people apply emotionally-mediated behaviors to both types of interlocutors. While there were small differences in magnitude by participant gender, the overall patterns were similar for women and men, supporting a nuanced picture of emotional vocal alignment.