Pruned RNN-T for fast, memory-efficient ASR training

Pruned RNN-T for fast, memory-efficient ASR training
复制标题

修剪 RNN-T 以实现快速、内存高效的 ASR 训练

DOI:
--
复制
发表时间:
2022
期刊:
Interspeech
影响因子:
--
通讯作者:
Daniel Povey
Daniel Povey
中科院分区:
--
文献类型:
--
作者:
Fangjun Kuang;Liyong Guo;Wei Kang;Long Lin;Mingshuang Luo;Zengwei Yao;Daniel Povey

文献摘要

被引文献

相似文献

用于语音识别的RNN-Transducer(RNN-T)框架越来越受欢迎,特别是对于部署的实时ASR系统,因为它将高准确性与自然流识别相结合。RNN-T的缺点之一是它的损失函数计算起来相对较慢,并且会占用大量内存。在词汇量很大的情况下,过多的GPU内存使用可能会使使用RNN-T损失变得不切实际:例如,对于基于中文字符的ASR。我们介绍了一种更快,更节省内存的RNN-T损失计算方法。我们首先使用一个简单的joiner网络获得RNN-T递归的修剪边界,该网络在编码器和解码器嵌入中是线性的;我们可以在不使用太多内存的情况下评估它。然后,我们使用这些修剪边界来评估完整的非线性joiner网络。
The RNN-Transducer (RNN-T) framework for speech recognition has been growing in popularity, particularly for deployed real-time ASR systems, because it combines high accuracy with naturally streaming recognition. One of the drawbacks of RNN-T is that its loss function is relatively slow to compute, and can use a lot of memory. Excessive GPU memory usage can make it impractical to use RNN-T loss in cases where the vocabulary size is large: for example, for Chinese character-based ASR. We introduce a method for faster and more memory-efficient RNN-T loss computation. We first obtain pruning bounds for the RNN-T recursion using a simple joiner network that is linear in the encoder and decoder embeddings; we can evaluate this without using much memory. We then use those pruning bounds to evaluate the full, non-linear joiner network.