Knowledge Distillation via Module Replacing for Automatic Speech Recognition with Recurrent Neural Network Transducer

Knowledge Distillation via Module Replacing for Automatic Speech Recognition with Recurrent Neural Network Transducer
复制标题

DOI:
10.21437/interspeech.2022-500
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Kaiqi Zhao;H. Nguyen;Animesh Jain;Nathan Susanj;A. Mouchtaris;Lokesh;Gupta;Ming Zhao
Kaiqi Zhao;H. Nguyen;Animesh Jain;Nathan Susanj;A. Mouchtaris;Lokesh;Gupta;Ming Zhao
中科院分区:
其他
文献类型:
--
作者:
Kaiqi Zhao;H. Nguyen;Animesh Jain;Nathan Susanj;A. Mouchtaris;Lokesh;Gupta;Ming Zhao

文献摘要

被引文献

相似文献

自动语音识别(ASR)越来越多地由智能虚拟助手等边缘应用程序使用。 。 lcmr替换(LCMR)评估是一种新型的非线性课程学习驱动的替代策略,通过通过LCMR的动态,平滑机制更新替代率,进一步提高了性能,学生和教师能够在梯度水平上进行交互,而Transser比常规KD更有效地了解。表明与常规KD相比,LCMR将单词误差(WER)降低14.47%-33.24%。
Automatic Speech Recognition (ASR) is increasingly used by edge applications such as intelligent virtual assistants. However, state-of-the-art ASR models such as Recurrent Neural Network - Transducer (RNN-T) are computationally intensive on resource-constrained edge devices. Knowledge Distillation (KD) is a promising approach to compress large models by us-ing a large model (”teacher”) to train a small model (”student”). This paper proposes a novel KD method called Log-Curriculum based Module Replacing (LCMR) for RNN-T. LCMR compresses RNN-T and addresses its unique characteristics by re-placing teacher modules including multiple LSTM/Dense layers with substitutional student modules that contain less Long Short Term Memory (LSTM)/Dense layers. LCMR employs a novel nonlinear Curriculum Learning driven replacement strategy to further improve the performance by updating replacing rates with a dynamic, smoothing mechanism. Under LCMR, the student and teacher are able to interact at gradient level, and tranfser knowledge more effectively than conventional KD. Evaluation shows that LCMR reduces word-error-rate (WER) by 14.47%-33.24% relative compared to conventional KD.