Long-Term, in-the-Wild Study of Feedback about Speech Intelligibility for K-12 Students Attending Class via a Telepresence Robot

Long-Term, in-the-Wild Study of Feedback about Speech Intelligibility for K-12 Students Attending Class via a Telepresence Robot
复制标题

通过远程呈现机器人上课的 K-12 学生语音清晰度反馈的长期野外研究

DOI:
10.1145/3462244.3479893
复制
发表时间:
2021
期刊:
Proceedings of the 2021 International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Maja J. Matari'c
Maja J. Matari'c
中科院分区:
--
文献类型:
--
作者:
Matthew Rueben;M. Syed;Emily London;Mark Camarena;Eunsook Shin;Yulun Zhang;Timothy S. Wang;Thomas R. Groechel;Rhianna Lee;Maja J. Matari'c

文献摘要

被引文献

相似文献

远程呈现机器人为远程用户提供临场感、体现力和移动性,使它们成为居家 K-12 学生的有希望的选择。然而,机器人操作员很难知道在偏远和嘈杂的教室环境中他们的声音有多好。一种解决方案是估计操作员对听众的语音清晰度,以便向操作员提供相关反馈。这项工作首次对在家远程上课的 K-12 学生的语音清晰度反馈系统进行了评估。在我们的四次长期野外部署中,我们发现学生以不同的音量说话,而不是调整机器人的音量,并且需要详细的音频校准和网络延迟反馈。我们还贡献了关于课堂听众向居家学生提供的多模式理解线索的类型和频率的第一个发现。通过对 700 多个提示进行注释和分类,我们发现最常见的提示模式是对话轮次时间和言语内容。总体而言,对话轮次线索出现的频率更高,而言语内容线索包含更多信息,可能是负面线索最常见的形式。我们的工作为远程呈现系统提供了建议,这些系统可以进行干预,以确保远程用户的声音被听到。
Telepresence robots offer presence, embodiment, and mobility to remote users, making them promising options for homebound K-12 students. It is difficult, however, for robot operators to know how well they are being heard in remote and noisy classroom environments. One solution is to estimate the operator’s speech intelligibility to their listeners in order to provide feedback about it to the operator. This work contributes the first evaluation of a speech intelligibility feedback system for homebound K-12 students attending class remotely. In our four long-term, in-the-wild deployments we found that students speak at different volumes instead of adjusting the robot’s volume, and that detailed audio calibration and network latency feedback are needed. We also contribute the first findings about the types and frequencies of multimodal comprehension cues given to homebound students by listeners in the classroom. By annotating and categorizing over 700 cues, we found that the most common cue modalities were conversation turn timing and verbal content. Conversation turn timing cues occurred more frequently overall, whereas verbal content cues contained more information and might be the most frequent modality for negative cues. Our work provides recommendations for telepresence systems that could intervene to ensure that remote users are being heard.