Massively Multilingual ASR: A Lifelong Learning Solution
Massively Multilingual ASR: A Lifelong Learning Solution
复制标题
大规模多语言 ASR:终身学习解决方案
DOI:
10.1109/icassp43922.2022.9746594
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Manasa Prasad
中科院分区:
文献类型:
--
作者:
Bo Li;Ruoming Pang;Yu Zhang;Tara N. Sainath;Trevor Strohman;Parisa Haghani;Yun Zhu;B. Farris;Neeraj Gaur;Manasa Prasad
The development of end-to-end models has largely sped up the research in massively multilingual automatic speech recognition (MMASR). Previous research has demonstrated the feasibility to build high quality MMASR models. In this work, we study the impact of adding more languages and propose a lifelong learning approach to build high quality MMASR systems. Experiments on a 66-language Voice Search task show that we can take a model built on 15 languages and continue training to obtain a 32-language model and similarly to further build a 67-language model. More importantly, models developed in this way achieve better quality compared to those trained from scratch. It maintains similar performance on old languages and achieves competitive results on new ones. This would potentially speed up the development of universal ASR models that recognize speech from any language, any domain and any environment by reusing knowledge learned beforehand.