Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings
Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings
复制标题
在资源匮乏的情况下通过反向翻译进行端到端 ASR 预训练
DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
S. Khudanpur
中科院分区:
文献类型:
--
作者:
Matthew Wiesner;Adithya Renduchintala;Shinji Watanabe;Chunxi Liu;N. Dehak;S. Khudanpur
We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because they depend on having sufficient target-side text to train the attention and decoder networks. In this paper we address this shortcoming by pretraining our network parameters using only text-based data and transcribed speech from other languages. We analyze the relative contributions of both sources of data. Across 3 test languages, our text-based approach resulted in a 20% average relative improvement over a text-based augmentation technique without pretraining. Using transcribed speech from nearby languages gives a further 20-30% relative reduction in character error rate.