Noise-Robust Key-Phrase Detectors for Automated Classroom Feedback

Noise-Robust Key-Phrase Detectors for Automated Classroom Feedback
复制标题

DOI:
10.1109/icassp40776.2020.9053173
复制
发表时间:
2020-05
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Brian Zylich;J. Whitehill
Brian Zylich;J. Whitehill
中科院分区:
其他
文献类型:
--
作者:
Brian Zylich;J. Whitehill

文献摘要

被引文献

相似文献

为了给老师提供关于课堂的自动反馈,我们研究了如何训练关键短语的自动语音检测器,比如干得好、谢谢、请、不客气。这种语言表达了老师对学生的支持和尊重,是已建立的CLASS[1]课堂观察协议中使用的行为标记之一。学校教室嘈杂且包含重叠的语音,即使采用最先进的方法,也为自动语音识别(ASR)提供了极具挑战性的环境。我们在一个中等规模但高度定制的课堂演讲数据集上使用分层多任务学习(MTL)训练深度神经网络。与两种最先进的通用语音识别ASR系统(b谷歌[2]和Deep-Speech[3])相比,我们的系统在匹配精度(30%)的同时,显著提高了召回率(50.4%对20.5%)。此外,我们的系统预测与CLASS的几个维度相关。
With the goal of giving teachers automated feedback about their classrooms, we investigate how to train automatic speech detectors of key phrases such as good job, thank you, please, and you’re welcome. This kind of language conveys support and respect from teacher to student and is one of the behavioral markers used in the established CLASS [1] classroom observation protocol. School classrooms are noisy and contain overlapping speech, presenting a highly challenging environment for automatic speech recognition (ASR), even for state-of-the-art approaches. We train deep neural networks using hierarchical multitask learning (MTL) on a modest-sized but highly-tailored dataset of classroom speech. Compared to 2 state-of-the-art ASR systems for general-purpose speech recognition (Google [2] and Deep-Speech [3]), our system delivers a substantially improved recall rate (50.4% versus 20.5%) while matching their precision (30%). Moreover, our system’s predictions correlate with several dimensions of the CLASS.