Massively Multilingual ASR: A Lifelong Learning Solution

Massively Multilingual ASR: A Lifelong Learning Solution
复制标题

大规模多语言 ASR:终身学习解决方案

DOI:
10.1109/icassp43922.2022.9746594
复制
发表时间:
2022
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Manasa Prasad
Manasa Prasad
中科院分区:
--
文献类型:
--
作者:
Bo Li;Ruoming Pang;Yu Zhang;Tara N. Sainath;Trevor Strohman;Parisa Haghani;Yun Zhu;B. Farris;Neeraj Gaur;Manasa Prasad

文献摘要

被引文献

相似文献

端到端模型的发展很大程度上加速了大规模多语言自动语音识别(MMASR)的研究。先前的研究已经证明了构建高质量 MMASR 模型的可行性。在这项工作中,我们研究了添加更多语言的影响,并提出了一种终身学习方法来构建高质量的 MMASR 系统。在 66 种语言的语音搜索任务上的实验表明,我们可以采用基于 15 种语言构建的模型并继续训练以获得 32 种语言的模型,并类似地进一步构建 67 种语言的模型。更重要的是,与从头开始训练的模型相比,以这种方式开发的模型具有更好的质量。它在旧语言上保持相似的性能,并在新语言上取得有竞争力的结果。这可能会加速通用 ASR 模型的开发,该模型可以通过重用预先学习的知识来识别来自任何语言、任何领域和任何环境的语音。
The development of end-to-end models has largely sped up the research in massively multilingual automatic speech recognition (MMASR). Previous research has demonstrated the feasibility to build high quality MMASR models. In this work, we study the impact of adding more languages and propose a lifelong learning approach to build high quality MMASR systems. Experiments on a 66-language Voice Search task show that we can take a model built on 15 languages and continue training to obtain a 32-language model and similarly to further build a 67-language model. More importantly, models developed in this way achieve better quality compared to those trained from scratch. It maintains similar performance on old languages and achieves competitive results on new ones. This would potentially speed up the development of universal ASR models that recognize speech from any language, any domain and any environment by reusing knowledge learned beforehand.