End-to-End Word-Level Disfluency Detection and Classification in Children’s Reading Assessment

End-to-End Word-Level Disfluency Detection and Classification in Children’s Reading Assessment
复制标题

DOI:
10.1109/icassp49357.2023.10095555
复制
发表时间:
2023-06
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Lavanya Venkatasubramaniam;Vishal Sunder;E. Fosler-Lussier
Lavanya Venkatasubramaniam;Vishal Sunder;E. Fosler-Lussier
中科院分区:
其他
文献类型:
--
作者:
Lavanya Venkatasubramaniam;Vishal Sunder;E. Fosler-Lussier

文献摘要

相似文献

儿童言语的不流畅检测和分类对于阅读技能的教学具有巨大的潜力。对儿童言语的单词级评估可以帮助教师有效地衡量学生的进步。因此,我们提出了一种新颖的基于注意力的模型,以完全端到端(E2E)的方式执行单词级不流畅检测和分类,使其快速且易于使用。我们开发了一种单词级不流利注释方案,用它来注释儿童阅读语音的数据集,即阅读竞赛数据集(READR)。我们还注释了现有 CMU Kids 语料库中的不流畅之处。所提出的模型在两个数据集上都显着优于使用强制对齐的传统级联基线。为了解决数据集中不可避免的类别不平衡问题,我们提出了一种名为 HiDeC(分层检测和分类)的新技术,该技术在 READR 和 CMU Kids 数据集上的相对 F1 分数分别提高了 23% 和 16% 的检测改进以及 3.8% 和 19.3% 的分类改进。
Disfluency detection and classification on children’s speech has a great potential for teaching reading skills. Word-level assessment of children’s speech can help teachers to effectively gauge their students’ progress. Hence, we propose a novel attention-based model to perform word-level disfluency detection and classification in a fully end-to-end (E2E) manner making it fast and easy to use. We develop a word-level disfluency annotation scheme using which we annotate a dataset of children read speech, the reading races dataset (READR). We also annotate disfluencies in the existing CMU Kids corpus. The proposed model significantly outperforms traditional cascaded baselines, which use forced alignments, on both datasets. To deal with the inevitable class-imbalance in the datasets, we propose a novel technique called HiDeC (Hierarchical Detection and Classification) which yields a detection improvement of 23% and 16% and a classification improvement of 3.8% and 19.3% relative F1-score on the READR and CMU Kids datasets respectively.