Multilingual Jointly Trained Acoustic and Written Word Embeddings

Multilingual Jointly Trained Acoustic and Written Word Embeddings
复制标题

DOI:
10.21437/interspeech.2020-2828
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
Yushi Hu;Shane Settle;Karen Livescu
Yushi Hu;Shane Settle;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
Yushi Hu;Shane Settle;Karen Livescu

文献摘要

相似文献

声学词嵌入(awe)是口语词段的向量表示。敬畏可以与字符序列的嵌入一起学习,以生成书面单词的语音有意义的嵌入,或基于声学的单词嵌入(AGWEs)。这种嵌入已被用于改进语音检索、识别和语音术语发现。在这项工作中,我们将这个想法扩展到多种低资源语言。我们联合训练AWE模型和AGWE模型,使用来自多种语言的语音转录数据。然后,预训练的模型可以用于看不见的零资源语言,或者对来自低资源语言的数据进行微调。我们还研究了独特的功能,作为电话标签的替代方案,以更好地共享跨语言信息。我们在12种语言的单词识别任务上测试了我们的模型。当对11种语言进行训练并对剩余的未见语言进行测试时,我们的模型优于传统的无监督方法,如动态时间翘曲。在一小时甚至十分钟的新语言数据上对预训练模型进行微调后,性能通常比仅使用目标语言数据进行训练要好得多。我们还发现语音监督提高了字符序列的性能,并且独特的特征监督有助于处理目标语言中看不见的电话。
Acoustic word embeddings (AWEs) are vector representations of spoken word segments. AWEs can be learned jointly with embeddings of character sequences, to generate phonetically meaningful embeddings of written words, or acoustically grounded word embeddings (AGWEs). Such embeddings have been used to improve speech retrieval, recognition, and spoken term discovery. In this work, we extend this idea to multiple low-resource languages. We jointly train an AWE model and an AGWE model, using phonetically transcribed data from multiple languages. The pre-trained models can then be used for unseen zero-resource languages, or fine-tuned on data from low-resource languages. We also investigate distinctive features, as an alternative to phone labels, to better share cross-lingual information. We test our models on word discrimination tasks for twelve languages. When trained on eleven languages and tested on the remaining unseen language, our model outperforms traditional unsupervised approaches like dynamic time warping. After fine-tuning the pre-trained models on one hour or even ten minutes of data from a new language, performance is typically much better than training on only the target-language data. We also find that phonetic supervision improves performance over character sequences, and that distinctive feature supervision is helpful in handling unseen phones in the target language.