On the Difficulties of Automatic Speech Recognition for Kindergarten-Aged Children

On the Difficulties of Automatic Speech Recognition for Kindergarten-Aged Children
复制标题

DOI:
10.21437/interspeech.2018-2297
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Gary Yeung;A. Alwan
Gary Yeung;A. Alwan
中科院分区:
其他
文献类型:
--
作者:
Gary Yeung;A. Alwan

文献摘要

相似文献

儿童自动语音识别(ASR)系统的性能落后于成人ASR。儿童ASR的确切问题和评价方法尚未得到充分的研究。机器人社区最近的工作表明,幼儿园语音的ASR尤其困难,尽管这个年龄段可能从基于语音的教育和诊断工具中贝内最多。我们的研究集中在特定年级(K-10)的ASR表现上,使用单词识别任务。评估了特定年级的ASR系统,特别关注对10岁以下儿童(5-6岁)的评估。实验包括使用特征空间最大似然线性回归(fMLLR),声道长度归一化(VTLN)和声门下共鸣(SGR)归一化的三音子模型研究等级特定的相互作用。我们的研究结果表明,幼儿园的ASR表现甚至比一年级的ASR要差得多,这可能是由于该年龄段的语音变异性很大。因此,ASR系统可能需要对幼儿园语音进行有针对性的评估,而不是以“儿童ASR”为幌子进行评估。此外,结果表明,在幼儿园语音匹配条件下训练的系统可能不如用一年级语音进行失配等级训练的系统合适。最后,我们分析了幼儿园ASR的语音偏误。
Automatic speech recognition (ASR) systems for children have lagged behind in performance when compared to adult ASR. The exact problems and evaluation methods for child ASR have not yet been fully investigated. Recent work from the robotics community suggests that ASR for kindergarten speech is especially difficult, even though this age group may benefit most from voice-based educational and diagnostic tools. Our study focused on ASR performance for specific grade levels (K-10) using a word identification task. Grade-specific ASR systems were evaluated, with particular attention placed on the evaluation of kindergarten-aged children (5-6 years old). Experiments included investigation of grade-specific interactions with triphone models using feature space maximum likelihood linear regression (fMLLR), vocal tract length normalization (VTLN), and subglottal resonance (SGR) normalization. Our results indicate that kindergarten ASR performs dramatically worse than even 1st grade ASR, likely due to large speech variability at that age. As such, ASR systems may require targeted evaluations on kindergarten speech rather than being evaluated under the guise of “child ASR.” Additionally, results show that systems trained in matched conditions on kindergarten speech may be less suitable than mismatched-grade training with 1st grade speech. Finally, we analyzed the phonetic errors made by the kindergarten ASR.