Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings

Pretraining by Backtranslation for End-to-End ASR in Low-Resource Settings
复制标题

在资源匮乏的情况下通过反向翻译进行端到端 ASR 预训练

DOI:
--
复制
发表时间:
2018
期刊:
Interspeech
影响因子:
--
通讯作者:
S. Khudanpur
S. Khudanpur
中科院分区:
--
文献类型:
--
作者:
Matthew Wiesner;Adithya Renduchintala;Shinji Watanabe;Chunxi Liu;N. Dehak;S. Khudanpur

文献摘要

被引文献

相似文献

我们探索在低资源环境中训练基于注意力的编码器-解码器ASR。这些模型在训练少量转录语音时表现不佳,部分原因是它们依赖于有足够的目标端文本来训练注意力和解码器网络。在本文中,我们通过仅使用基于文本的数据和其他语言的转录语音来预训练我们的网络参数来解决这个缺点。我们分析了这两种数据来源的相对贡献。在3种测试语言中,我们基于文本的方法比没有预训练的基于文本的增强技术平均相对提高了20%。使用附近语言的转录语音可以进一步降低20-30%的字符错误率。
We explore training attention-based encoder-decoder ASR in low-resource settings. These models perform poorly when trained on small amounts of transcribed speech, in part because they depend on having sufficient target-side text to train the attention and decoder networks. In this paper we address this shortcoming by pretraining our network parameters using only text-based data and transcribed speech from other languages. We analyze the relative contributions of both sources of data. Across 3 test languages, our text-based approach resulted in a 20% average relative improvement over a text-based augmentation technique without pretraining. Using transcribed speech from nearby languages gives a further 20-30% relative reduction in character error rate.