Deciphering Speech: a Zero-Resource Approach to Cross-Lingual Transfer in ASR

Deciphering Speech: a Zero-Resource Approach to Cross-Lingual Transfer in ASR
复制标题

DOI:
10.21437/interspeech.2022-10170
复制
发表时间:
2021-11
期刊:
--
影响因子:
--
通讯作者:
Ondrej Klejch;E. Wallington;P. Bell
Ondrej Klejch;E. Wallington;P. Bell
中科院分区:
其他
文献类型:
--
作者:
Ondrej Klejch;E. Wallington;P. Bell

文献摘要

被引文献

相似文献

我们提出了一种跨语言训练ASR系统的方法,使用绝对没有转录的训练数据从目标语言,没有语音知识的语言问题。我们的方法使用了一种新的应用程序的解密算法,该算法只考虑来自目标语言的未配对的语音和文本数据。我们将此解密应用于由在语言外语音语料库上训练的通用电话识别器生成的电话序列,然后进行平坦开始的半监督训练,以获得新语言的声学模型。据我们所知,这是第一个实现零资源跨语言ASR的实用方法,它不依赖于任何手工制作的语音信息。我们从GlobalPhone语料库中读取语音进行实验,并表明可以从目标语言中学习仅20分钟的数据的解密模型。当用于生成半监督训练的伪标签时,我们获得的WER范围从32.5%到1.9%,绝对差于在相同数据上训练的等效全监督模型。
We present a method for cross-lingual training an ASR system using absolutely no transcribed training data from the target language, and with no phonetic knowledge of the language in question. Our approach uses a novel application of a decipherment algorithm, which operates given only unpaired speech and text data from the target language. We apply this decipherment to phone sequences generated by a universal phone recogniser trained on out-of-language speech corpora, which we follow with flat-start semi-supervised training to obtain an acoustic model for the new language. To the best of our knowledge, this is the first practical approach to zero-resource cross-lingual ASR which does not rely on any hand-crafted phonetic information. We carry out experiments on read speech from the GlobalPhone corpus, and show that it is possible to learn a decipherment model on just 20 minutes of data from the target language. When used to generate pseudo-labels for semi-supervised training, we obtain WERs that range from 32.5% to just 1.9% absolute worse than the equivalent fully supervised models trained on the same data.