Unsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word Embeddings

Unsupervised Word Segmentation and Lexicon Discovery Using Acoustic Word Embeddings
复制标题

DOI:
10.1109/taslp.2016.2517567
复制
发表时间:
2016-03
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
H. Kamper;A. Jansen;S. Goldwater
中科院分区:
其他
文献类型:
--
作者:
H. Kamper;A. Jansen;S. Goldwater

文献摘要

被引文献

相似文献

在只有未标记的语音数据可用的环境中,语音技术需要在没有翻译、发音词典或语言建模文本的情况下开发。在为婴儿语言习得建模时也面临着类似的问题。在这些情况下,需要直接从语音音频中发现类别语言结构。我们提出了一种新的无监督贝叶斯模型,分割未标记的语音和集群的段到假设的单词组。结果是根据发现的词类型对输入语音进行完全无监督的标记化。在我们的方法中,一个潜在的词段(任意长度)嵌入在一个固定维的声学向量空间。该模型作为Gibbs采样器实现,然后在该空间中构建全词声学模型,同时联合执行分割。我们报告的单词错误率在一个小词汇量连接数字识别任务的映射无监督解码输出地面真理transmittance。该模型实现了约20%的错误率,比以前基于HMM的系统绝对性能高出约10%。此外,与基线相比,我们的模型不需要预先指定的词汇量。
In settings where only unlabeled speech data is available, speech technology needs to be developed without transcriptions, pronunciation dictionaries, or language modelling text. A similar problem is faced when modeling infant language acquisition. In these cases, categorical linguistic structure needs to be discovered directly from speech audio. We present a novel unsupervised Bayesian model that segments unlabeled speech and clusters the segments into hypothesized word groupings. The result is a complete unsupervised tokenization of the input speech in terms of discovered word types. In our approach, a potential word segment (of arbitrary length) is embedded in a fixed-dimensional acoustic vector space. The model, implemented as a Gibbs sampler, then builds a whole-word acoustic model in this space while jointly performing segmentation. We report word error rates in a small-vocabulary connected digit recognition task by mapping the unsupervised decoded output to ground truth transcriptions. The model achieves around 20% error rate, outperforming a previous HMM-based system by about 10% absolute. Moreover, in contrast to the baseline, our model does not require a pre-specified vocabulary size.