A Correspondence Variational Autoencoder for Unsupervised Acoustic Word Embeddings

A Correspondence Variational Autoencoder for Unsupervised Acoustic Word Embeddings
复制标题

DOI:
--
复制
发表时间:
2020-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Puyuan Peng;H. Kamper;Karen Livescu
Puyuan Peng;H. Kamper;Karen Livescu
中科院分区:
其他
文献类型:
--
作者:
Puyuan Peng;H. Kamper;Karen Livescu

文献摘要

被引文献

相似文献

我们提出了一个新的无监督模型,用于将可变持续时间的语音段映射到固定维度的表示。由此产生的声学词嵌入可以形成低资源和零资源语言的搜索、发现和索引系统的基础。我们的模型,我们称之为最大采样对应变分自编码器(MCVAE),是一个递归神经网络(RNN),使用一种新的自监督对应丢失进行训练,鼓励同一单词的不同实例的嵌入之间的一致性。我们的训练方案通过使用和比较来自近似后验分布的多个样本来改进以前的对应训练方法。在零资源设置中,MCVAE可以通过使用经由无监督术语发现系统发现的类词片段以无监督方式训练,而不需要任何地面实况词对。在这种设置和半监督低资源设置(具有有限的地面实况词对)中,MCVAE的性能优于以前的最先进模型,例如基于Siamese,CAE和VAE的RNN。
We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis of search, discovery, and indexing systems for low- and zero-resource languages. Our model, which we refer to as a maximal sampling correspondence variational autoencoder (MCVAE), is a recurrent neural network (RNN) trained with a novel self-supervised correspondence loss that encourages consistency between embeddings of different instances of the same word. Our training scheme improves on previous correspondence training approaches through the use and comparison of multiple samples from the approximate posterior distribution. In the zero-resource setting, the MCVAE can be trained in an unsupervised way, without any ground-truth word pairs, by using the word-like segments discovered via an unsupervised term discovery system. In both this setting and a semi-supervised low-resource setting (with a limited set of ground-truth word pairs), the MCVAE outperforms previous state-of-the-art models, such as Siamese-, CAE- and VAE-based RNNs.