Linear dynamic models for automatic speech recognition

Linear dynamic models for automatic speech recognition
复制标题

DOI:
--
复制
发表时间:
2004-06
期刊:
--
影响因子:
--
通讯作者:
Joe Frankel
Joe Frankel
中科院分区:
其他
文献类型:
--
作者:
Joe Frankel

文献摘要

被引文献

相似文献

大多数自动语音识别(ASR)系统依赖于隐马尔可夫模型(HMM),其中与每个状态相关的输出分布由对角协方差高斯的混合建模。通常通过将时间导数附加到特征向量来包含动态信息。这种方法虽然成功,但做出了增强特征向量的帧独立性的错误假设,并且忽略了参数化语音信号中的空间相关性。本论文旨在通过应用状态空间模型形式(线性动态模型(LDM))探索 ASR 声学建模来解决这些缺点。 LDM 不是对单个数据帧进行建模,而是对整个语音片段进行表征。通过连续空间的自回归状态演化给出了潜在动力学的马尔可夫模型,并且特征维度之间的空间相关性被吸收到观察过程的结构中。 LDM 之前已应用于语音识别,但是使用了平滑的高斯-马尔可夫形式,忽略了子空间建模的潜力。连续的动态状态意味着信息沿着每个段的长度传递。此外,如果允许状态跨段边界连续,则系统中会内置长程依赖性,并且连续段的独立性假设会被放松。状态提供了时间相关性的显式模型,该模型将这种方法与基于帧和一些基于段的模型区分开来,在这些模型中数据的排序并不重要。这种模型的好处在细分市场内部和细分市场之间进行了检验。 LDM 非常适合对平滑变化、连续但有噪声的轨迹进行建模,例如在测量的关节数据中发现的轨迹。使用 MOCHA 语料库中与说话者相关的数据,对声学、发音和组合声学发音特征进行建模的系统的性能进行了比较。除了测量的发音参数之外,实验还使用经过训练的神经网络的输出来执行发音反转映射。与说话人无关的 TIMIT 语料库为更大规模的纯声学实验提供了基础。分类任务提供了一种理想的方法来比较建模选择,而不会受到识别搜索错误的混杂影响,并用于探索状态维度的选择、前端声学参数化和参数初始化等问题。分段模型的识别通常比基于帧的模型的计算成本更高。与帧级模型不同,并不总是可以共享发生在具有不同开始和结束时间的假设片段内的观察序列的似然计算。此外,维特比准则不一定适用于帧级别。这项工作以带有 A* 搜索的堆栈解码器的形式引入了一种新颖的分段模型解码方法。这种方案允许灵活选择声学和语言模型,因为维特比准则不是搜索的组成部分,并且假设生成独立于特定的语言模型。此外,搜索的时间异步排序意味着仅扩展可能的路径,因此评估最少数量的模型。解码器用于给出源自 MOCHA 和 TIMIT 语料库的特征集的完整识别结果。使用传统的训练/测试划分和语言模型的选择,以便可以将结果直接与其他研究中的结果进行比较。解码器还用于实现维特比训练,模型参数交替更新,然后用于重新对齐训练数据。
The majority of automatic speech recognition (ASR) systems rely on hidden Markov models (HMM), in which the output distribution associated with each state is modelled by a mixture of diagonal covariance Gaussians. Dynamic information is typically included by appending time-derivatives to feature vectors. This approach, whilst successful, makes the false assumption of framewise independence of the augmented feature vectors and ignores the spatial correlations in the parametrised speech signal. This dissertation seeks to address these shortcomings by exploring acoustic modelling for ASR with an application of a form of state-space model, the linear dynamic model (LDM). Rather than modelling individual frames of data, LDMs characterise entire segments of speech. An auto-regressive state evolution through a continuous space gives a Markovian model of the underlying dynamics, and spatial correlations between feature dimensions are absorbed into the structure of the observation process. LDMs have been applied to speech recognition before, however a smoothed Gauss-Markov form was used which ignored the potential for subspace modelling. The continuous dynamical state means that information is passed along the length of each segment. Furthermore, if the state is allowed to be continuous across segment boundaries, long range dependencies are built into the system and the assumption of independence of successive segments is loosened. The state provides an explicit model of temporal correlation which sets this approach apart from frame-based and some segment-based models where the ordering of the data is unimportant. The benefits of such a model are examined both within and between segments. LDMs are well suited to modelling smoothly varying, continuous, yet noisy trajectories such as found in measured articulatory data. Using speaker-dependent data from the MOCHA corpus, the performance of systems which model acoustic, articulatory, and combined acoustic-articulatory features are compared. As well as measured articulatory parameters, experiments use the output of neural networks trained to perform an articulatory inversion mapping. The speaker-independent TIMIT corpus provides the basis for larger scale acoustic-only experiments. Classification tasks provide an ideal means to compare modelling choices without the confounding influence of recognition search errors, and are used to explore issues such as choice of state dimension, front-end acoustic parametrisation and parameter initialisation. Recognition for segment models is typically more computationally expensive than for frame-based models. Unlike frame-level models, it is not always possible to share likelihood calculations for observation sequences which occur within hypothesised segments that have different start and end times. Furthermore, the Viterbi criterion is not necessarily applicable at the frame level. This work introduces a novel approach to decoding for segment models in the form of a stack decoder with A∗ search. Such a scheme allows flexibility in the choice of acoustic and language models since the Viterbi criterion is not integral to the search, and hypothesis generation is independent of the particular language model. Furthermore, the time-asynchronous ordering of the search means that only likely paths are extended, and so a minimum number of models are evaluated. The decoder is used to give full recognition results for feature-sets derived from the MOCHA and TIMIT corpora. Conventional train/test divisions and choice of language model are used so that results can be directly compared to those in other studies. The decoder is also used to implement Viterbi training, in which model parameters are alternately updated and then used to re-align the training data.