Accessibility Evaluation of Classroom Captions

Accessibility Evaluation of Classroom Captions
复制标题

课堂字幕的无障碍评估

DOI:
10.1145/2543578
复制
发表时间:
2014
期刊:
ACM Transactions on Accessible Computing (TACCESS)
影响因子:
--
通讯作者:
Jeffrey P. Bigham
Jeffrey P. Bigham
中科院分区:
--
文献类型:
--
作者:
R. Kushalnagar;Walter S. Lasecki;Jeffrey P. Bigham

文献摘要

被引文献

相似文献

实时字幕使聋人和听力困难(DHH)的人能够通过将其转换为视觉文本,以不到5秒的延迟来跟随课堂讲座和其他听觉演讲。保持较短的延迟可以让最终用户跟踪并参与对话。本文重点讨论导致实时字幕难以实现的基本问题:顺序键盘输入比说话慢得多。我们首先调查了YouTube上240个一小时长的带标题讲座的音频特征,例如说话的速度和持续时间。然后,我们分析了这些特征如何影响字幕生成和可读性,特别是考虑到我们的人力协作字幕方法。我们注意到,这些特征中的大多数也存在于更一般的领域。对于我们的字幕比较评估,我们使用所有三种字幕方法实时转录课堂讲座。我们招募了48名参与者(24 DHH)在眼动追踪实验室中观看这些课堂成绩单。我们以随机、平衡的顺序呈现这些标题。我们发现,听力和DHH的参与者更喜欢和遵循的协作字幕比自动语音识别(ASR)或专业人员产生的字幕,由于更一致的流动所产生的字幕。这些结果表明,即使在速度突然爆发的情况下,也有可能可靠地捕获语音,以及生成“增强”字幕,这与其他人力字幕方法不同。
Real-time captioning enables deaf and hard of hearing (DHH) people to follow classroom lectures and other aural speech by converting it into visual text with less than a five second delay. Keeping the delay short allows end-users to follow and participate in conversations. This article focuses on the fundamental problem that makes real-time captioning difficult: sequential keyboard typing is much slower than speaking. We first surveyed the audio characteristics of 240 one-hour-long captioned lectures on YouTube, such as speed and duration of speaking bursts. We then analyzed how these characteristics impact caption generation and readability, considering specifically our human-powered collaborative captioning approach. We note that most of these characteristics are also present in more general domains. For our caption comparison evaluation, we transcribed a classroom lecture in real-time using all three captioning approaches. We recruited 48 participants (24 DHH) to watch these classroom transcripts in an eye-tracking laboratory. We presented these captions in a randomized, balanced order. We show that both hearing and DHH participants preferred and followed collaborative captions better than those generated by automatic speech recognition (ASR) or professionals due to the more consistent flow of the resulting captions. These results show the potential to reliably capture speech even during sudden bursts of speed, as well as for generating “enhanced” captions, unlike other human-powered captioning approaches.