A self-transcribing speech corpus: collecting continuous speech with an online educational game

A self-transcribing speech corpus: collecting continuous speech with an online educational game
复制标题

自转录语音语料库:通过在线教育游戏收集连续语音

DOI:
--
复制
发表时间:
2009
期刊:
Slate
影响因子:
--
通讯作者:
A. M. Sutherland
A. M. Sutherland
中科院分区:
--
文献类型:
--
作者:
A. Gruenstein;Ian McGraw;A. M. Sutherland

文献摘要

被引文献

相似文献

我们描述了一种通过使用名为“Voice Scatter”的在线教育游戏来收集拼字转录的连续语音数据的新颖方法,在该游戏中,玩家通过使用语音将术语与其定义进行匹配来学习抽认卡。我们分析了 Voice Scatter 公开发布的前 22 天内收集的包含 30,938 条话语、总计 27.63 小时语音的语料库。尽管每个单独的游戏仅涵盖很小的词汇量,但语料库中的语音识别假设总共包含 21,758 个不同的单词。我们证明,Amazon Mechanical Turk 可用于快速、廉价地以拼字法转录语料库中的话语,并且具有接近专家的准确性。此外,我们提出了一种过滤技术,可以自动识别 39% 数据的子语料库,其中识别假设可以被视为人类质量的转录本。我们证明了这种自转录数据对于声学模型适应的有用性。
We describe a novel approach to collecting orthographically transcribed continuous speech data through the use of an online educational game called Voice Scatter, in which players study flashcards by using speech to match terms with their definitions. We analyze a corpus of 30,938 utterances, totaling 27.63 hours of speech, collected during the first 22 days that Voice Scatter was publicly available. Though each individual game covers only a small vocabulary, in aggregate speech recognition hypotheses in the corpus contain 21,758 distinct words. We show that Amazon Mechanical Turk can be used to orthographically transcribe utterances in the corpus quickly and cheaply, with near-expert accuracy. Moreover, we present a filtering technique that automatically identifies a sub-corpus of 39% of the data for which recognition hypotheses can be considered human-quality transcripts. We demonstrate the usefulness of such self-transcribed data for acoustic model adaptation.