Acoustic-phonetic properties of Siri- and human-directed speech

Acoustic-phonetic properties of Siri- and human-directed speech
复制标题

DOI:
10.1016/j.wocn.2021.101123
复制
发表时间:
2021-12-20
影响因子:
1.9
通讯作者:
Zellou, Georgia
Zellou, Georgia
中科院分区:
人文科学1区
文献类型:
--
作者:
Cohn, Michelle;Segedin, Bruno Ferenc;Zellou, Georgia

文献摘要

被引文献

相似文献

数百万人在日常生活中与语音激活的人工智能 (voice-AI) 系统进行语音交互。这项研究探讨了说话者是否具有特定于语音人工智能的语域(相对于他们对成年人的言语)。此外,这项研究还测试了说话者是否针对语音人工智能和人类对话者制定了有针对性的纠错策略。在一项使用预先录制的 Siri 和人声的伪交互式任务中,参与者在句子中产生目标词。在每一轮中,在对话者初步产生和反馈之后,参与者以三种响应类型之一重复该句子:在正确的单词识别之后,尾声错误,或对话者犯的元音错误。在两项研究中,两位对话者的理解错误率各不相同(错误率较低与较高)。发现了语域差异:参与者说话声音更大,平均 f0 更低,Siri-DS 中的 f0 范围更小。 Siri-DS 中的许多差异是随着交互过程中的动态调整而出现的。此外,错误率决定了寄存器差异的实现方式。观察到一种有针对性的错误纠正:扬声器在 Siri-DS 的尾声修复中产生更多元音过度发音。总而言之,这些发现有助于我们理解语音语域和说话者与对话者互动的动态本质。
Millions of people engage in spoken interactions with voice activated artificially intelligent (voice-AI) systems in their everyday lives. This study explores whether speakers have a voice-AI-specific register, relative to their speech toward an adult human. Furthermore, this study tests if speakers have targeted error correction strategies for voice-AI and human interlocutors. In a pseudo-interactive task with pre-recorded Siri and human voices, participants produced target words in sentences. In each turn, following an initial production and feedback from the interlocutor, participants repeated the sentence in one of three response types: after correct word identification, a coda error, or a vowel error made by the interlocutor. Across two studies, the rate of comprehension errors made by both interlocutors was varied (lower vs. higher error rate). Register differences are found: participants speak louder, with a lower mean f0, and with a smaller f0 range in Siri-DS. Many differences in Siri-DS emerged as dynamic adjustments over the course of the interaction. Additionally, error rate shapes how register differences are realized. One targeted error correction was observed: speakers produce more vowel hyperarticulation in coda repairs in Siri-DS. Taken together, these findings contribute to our understanding of speech register and the dynamic nature of talker-interlocutor interactions.