A hierarchical system for word discovery exploiting DTW-based initialization

A hierarchical system for word discovery exploiting DTW-based initialization
复制标题

利用基于 DTW 的初始化进行单词发现的分层系统

DOI:
10.1109/asru.2013.6707761
复制
发表时间:
2013
期刊:
2013 IEEE Workshop on Automatic Speech Recognition and Understanding
影响因子:
--
通讯作者:
Haeb-Umbach
Haeb-Umbach
中科院分区:
--
文献类型:
--
作者:
Walter;Korthals;Haeb-Umbach

文献摘要

参考文献

被引文献

相似文献

仅从口语输入中发现语言的语言结构需要两个步骤:语音发现和词汇发现。第一个是关于确定的分类子词单位库存和相关的基本声学,而第二个目的是在发现字作为重复模式的子词单位。这里提出的分层方法占在第一阶段的分类错误,通过建模的发音一个字的子字单元的概率:一个隐马尔可夫模型与离散的发射概率,发射所观察到的子字单元序列。我们描述了如何从语音输入中以完全无监督的方式学习系统。为了改进单词发音的训练的初始化,使用基于动态时间规整的声学模式发现系统的输出,因为它能够在输入数据中发现类似的时间序列。这种改进的初始化,只使用弱监督,导致数字识别任务的字错误率降低了40%。
Discovering the linguistic structure of a language solely from spoken input asks for two steps: phonetic and lexical discovery. The first is concerned with identifying the categorical subword unit inventory and relating it to the underlying acoustics, while the second aims at discovering words as repeated patterns of subword units. The hierarchical approach presented here accounts for classification errors in the first stage by modelling the pronunciation of a word in terms of subword units probabilistically: a hidden Markov model with discrete emission probabilities, emitting the observed subword unit sequences. We describe how the system can be learned in a completely unsupervised fashion from spoken input. To improve the initialization of the training of the word pronunciations, the output of a dynamic time warping based acoustic pattern discovery system is used, as it is able to discover similar temporal sequences in the input data. This improved initialization, using only weak supervision, has led to a 40% reduction in word error rate on a digit recognition task.
用于音频语义分析的无监督结构发现
DOI: --
发表时间: 2012
期刊: Neural Information Processing Systems
影响因子: --
作者:
Sourish Chaudhuri;B. Raj
通讯作者: B. Raj