Improved Acoustic Word Embeddings for Zero-Resource Languages Using Multilingual Transfer

Improved Acoustic Word Embeddings for Zero-Resource Languages Using Multilingual Transfer
复制标题

DOI:
10.1109/taslp.2021.3060805
复制
发表时间:
2020-06
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Kamper;Yevgen Matusevych;S. Goldwater
H. Kamper;Yevgen Matusevych;S. Goldwater
中科院分区:
其他
文献类型:
--
作者:
H. Kamper;Yevgen Matusevych;S. Goldwater

文献摘要

相似文献

声学词嵌入是可变长度语音片段的固定维度表示。当传统语音识别无法实现时,这种嵌入可以构成语音搜索、索引和发现系统的基础。在零资源设置中,未标记的语音是唯一可用的资源,我们需要一种能够在任意语言上提供鲁棒嵌入的方法。在这里,我们探索多语言迁移:我们在来自多种资源丰富的语言的标记数据上训练单个监督嵌入模型,然后将其应用于看不见的零资源语言。我们考虑三种多语言递归神经网络(RNN)模型:在所有训练语言的联合词汇上训练的分类器;训练有素的连体 RNN 来区分多种语言中相同和不同的单词;以及经过训练以重建单词对的对应自动编码器 (CAE) RNN。在六种目标语言的单词辨别任务中,所有这些模型都优于在零资源语言本身上训练的最先进的无监督模型,平均精度相对提高了 30% 以上。当仅使用几种训练语言时,多语言 CAE-RNN 表现更好,但使用更多训练语言时,其他多语言模型表现相似。使用更多的训练语言通常是有益的,但对某些语言的改进是微乎其微的。我们提出的探索实验表明,CAE-RNN 比其他多语言模型编码更多的语音、单词持续时间、语言身份和说话人信息。
Acoustic word embeddings are fixed-dimensional representations of variable-length speech segments. Such embeddings can form the basis for speech search, indexing and discovery systems when conventional speech recognition is not possible. In zero-resource settings where unlabelled speech is the only available resource, we need a method that gives robust embeddings on an arbitrary language. Here we explore multilingual transfer: we train a single supervised embedding model on labelled data from multiple well-resourced languages and then apply it to unseen zero-resource languages. We consider three multilingual recurrent neural network (RNN) models: a classifier trained on the joint vocabularies of all training languages; a Siamese RNN trained to discriminate between same and different words from multiple languages; and a correspondence autoencoder (CAE) RNN trained to reconstruct word pairs. In a word discrimination task on six target languages, all of these models outperform state-of-the-art unsupervised models trained on the zero-resource languages themselves, giving relative improvements of more than 30% in average precision. When using only a few training languages, the multilingual CAE-RNN performs better, but with more training languages the other multilingual models perform similarly. Using more training languages is generally beneficial, but improvements are marginal on some languages. We present probing experiments which show that the CAE-RNN encodes more phonetic, word duration, language identity and speaker information than the other multilingual models.