Fully unsupervised small-vocabulary speech recognition using a segmental Bayesian model

Fully unsupervised small-vocabulary speech recognition using a segmental Bayesian model
复制标题

DOI:
10.21437/interspeech.2015-239
复制
发表时间:
2015
期刊:
--
影响因子:
--
通讯作者:
H. Kamper;A. Jansen;S. Goldwater
H. Kamper;A. Jansen;S. Goldwater
中科院分区:
其他
文献类型:
--
作者:
H. Kamper;A. Jansen;S. Goldwater

文献摘要

相似文献

当前有监督的语音技术在很大程度上依赖于抄录的语音和发音词典。在单独使用无标记的语音数据的设置中,需要无监督的方法直接从音频发现分类语言结构。我们提出了一种新颖的贝叶斯模型,该模型将未标记的输入语音分为单词式单元,从而使语音完全无监督的转录以发现的单词类型的方式。在我们的方法中,将(任意长度)的潜在单词段嵌入在固定维空间中。然后,该模型(以Gibbs采样器的形式实现)然后在该空间中构建一个全词的声学模型,同时共同进行分割。我们通过将无监督的输出映射到地面真相转录中,以在连接的数字识别任务中报告单词错误率。我们的模型的表现优于先前开发的基于HMM的系统,即使模型不受限制地仅发现数据中存在的11个单词类型。索引术语:无监督的语音处理,单词发现,语音细分,无监督的学习,分段模型
Current supervised speech technology relies heavily on transcribed speech and pronunciation dictionaries. In settings where unlabelled speech data alone is available, unsupervised methods are required to discover categorical linguistic structure directly from the audio. We present a novel Bayesian model which segments unlabelled input speech into word-like units, resulting in a complete unsupervised transcription of the speech in terms of discovered word types. In our approach, a potential word segment (of arbitrary length) is embedded in a fixed-dimensional space; the model (implemented as a Gibbs sampler) then builds a whole-word acoustic model in this space while jointly doing segmentation. We report word error rates in a connected digit recognition task by mapping the unsupervised output to ground truth transcriptions. Our model outperforms a previously developed HMM-based system, even when the model is not constrained to discover only the 11 word types present in the data. Index Terms: unsupervised speech processing, word discovery, speech segmentation, unsupervised learning, segmental models