THE OGI KIDS’ SPEECH CORPUS AND RECOGNIZERS

THE OGI KIDS’ SPEECH CORPUS AND RECOGNIZERS
复制标题

OGI 儿童语音语料库和识别器

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
R. Cole
R. Cole
中科院分区:
--
文献类型:
--
作者:
Khaldoun Shobaki;John;R. Cole

文献摘要

被引文献

相似文献

我们描述了一个语料库的儿童语音,称为OGI儿童语音语料库,和扬声器和词汇独立的识别系统训练和评估这些数据。语料库由1100名幼儿园到10年级的儿童的提示和自发语音组成。提示的语音被呈现为出现在动画人物(Baldi)下方的文本,该动画人物产生与记录的提示同步的准确可见的语音。语音和文本由孤立的单词、句子和数字串组成。语音识别器的训练使用HMM/ANN框架,训练数据取自与语音段在语料库中的孤立的话。语音段使用自动语音对齐。为了找出识别器能够推广到训练集中没有发现的新词的程度,我们进行了两个测试集评估:一个使用来自孤立说出的205个单词的一组新的话语(类似于用于训练识别器的数据),另一个使用来自提示句子的单词。结果是显着不同的(97.5%的孤立与37.9%的句子中的单词),我们探索的方法,可用于提高识别器的能力,概括为新词。
We describe a corpus of children’s speech, called the OGI Kids’ Speech corpus, and a speaker- and vocabulary-independent recognition system trained and evaluated with these data. The corpus is composed of both prompted and spontaneous speech from 1100 children from kindergarten through grade 10. The prompted speech was presented as text appearing below an animated character (Baldi) that produced accurate visible speech synchronized with recorded prompts. The speech and text consists of isolated words, sentences, and digit strings. A phonetic recognizer was trained using an HMM/ANN framework, with training data taken from intervals of speech associated with phonetic segments in the isolated words in the corpus. Phonetic segments were derived using automatic phonetic alignment. To find out how well the recognizer is able to generalize to new words not found in the training set, we performed two test-set evaluations: one using a new set of utterances from the set of 205 words spoken in isolation (similar to the data used to train the recognizer) and one using words from the prompted sentences. Results were dramatically different (97.5% for isolated vs. 37.9% for words in sentences), and we explore methods that may be used to improve the recognizer’s ability to generalize to new words.