BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking

BeParrot: Efficient Interface for Transcribing Unclear Speech via Respeaking
复制标题

BeParrot:通过复述转录不清楚语音的高效界面

DOI:
10.1145/3490099.3511164
复制
发表时间:
2022
期刊:
Proceedings of the 27th ACM International Conference on Intelligent User Interface
影响因子:
--
通讯作者:
and Masataka Goto
and Masataka Goto
中科院分区:
--
文献类型:
--
作者:
Riku Arakawa;Hiromu Yakura (equal contribution);and Masataka Goto

文献摘要

相似文献

将语音从音频文件转录为文本不仅对于以文本形式探索音频内容而且对于利用转录的数据作为源来训练语音模型(诸如自动语音识别(ASR)模型)是重要的任务。一个后校正的方法已被频繁地采用,以减少转录的时间成本,其中用户编辑的ASR模型的识别结果中的错误。然而,这种方法假设清晰的语音,而不是为不清晰的语音(例如具有高水平噪声或混响的语音)而设计的,这严重降低了ASR的准确性,并且需要许多手动校正。为了构建一种替代方法来转录不清楚的语音,我们引入了重新峰值化的想法,它主要用于在真实的时间内为电视节目创建字幕。在重说话中,熟练的人类重说话者重复所听到的语音作为阴影,并且他们的话语由ASR模型识别。虽然这种方法可以有效地转录不清楚的讲话,一个问题是,重新说话是一个高度认知要求的任务,往往需要广泛的培训,成为一个重新说话。我们使用BeParrot解决了这一问题,BeParrot是第一个专为重新发音而设计的界面,它允许新手用户通过两个关键功能从重新发音中受益,而无需进行大量培训:参数调整和发音反馈。我们的用户研究涉及60名群众工作者,表明他们可以用BeParrot转录不同类型的不清晰语音,比传统方法快32.2%,而不会损失转录的准确性。此外,工作人员的评论支持调整和反馈功能的设计,表现出继续使用BeParrot进行转录任务的意愿。我们的工作展示了我们如何利用机器学习技术的最新进展,在人在环方法的帮助下克服计算机本身仍然具有挑战性的领域。
Transcribing speech from audio files to text is an important task not only for exploring the audio content in text form but also for utilizing the transcribed data as a source to train speech models, such as automated speech recognition (ASR) models. A post-correction approach has been frequently employed to reduce the time cost of transcription where users edit errors in the recognition results of ASR models. However, this approach assumes clear speech and is not designed for unclear speech (such as speech with high levels of noise or reverberation), which severely degrades the accuracy of ASR and requires many manual corrections. To construct an alternative approach to transcribe unclear speech, we introduce the idea of respeaking, which has primarily been used to create captions for television programs in real time. In respeaking, a proficient human respeaker repeats the heard speech as shadowing, and their utterances are recognized by an ASR model. While this approach can be effective for transcribing unclear speech, one problem is that respeaking is a highly cognitively demanding task and extensive training is often required to become a respeaker. We address this point with BeParrot, the first interface designed for respeaking that allows novice users to benefit from respeaking without extensive training through two key features: parameter adjustment and pronunciation feedback. Our user study involving 60 crowd workers demonstrated that they could transcribe different types of unclear speech 32.2 % faster with BeParrot than with a conventional approach without losing the accuracy of transcriptions. In addition, comments from the workers supported the design of the adjustment and feedback features, exhibiting a willingness to continue using BeParrot for transcription tasks. Our work demonstrates how we can leverage recent advances in machine learning techniques to overcome the area that is still challenging for computers themselves with the help of a human-in-the-loop approach.