Deep neural network training for whispered speech recognition using small databases and generative model sampling

Deep neural network training for whispered speech recognition using small databases and generative model sampling
复制标题

DOI:
10.1007/s10772-017-9461-x
复制
发表时间:
2017-12-01
影响因子:
--
通讯作者:
Hansen, John H. L.
Hansen, John H. L.
中科院分区:
其他
文献类型:
--
作者:
Ghaffarzadegan, Shabnam;Boril, Hynek;Hansen, John H. L.

文献摘要

被引文献

相似文献

最先进的语音识别解决方案目前采用隐马尔可夫模型(HMM)来捕获语音信号中的时间变化性,并采用深度神经网络(DNN)来对HMM状态分布进行建模。在许多应用中,DNN-HMM混合系统的性能优于传统的HMM和高斯混合模型(GMM)混合系统。这种改进主要归功于DNN对更复杂数据结构建模的能力。然而,拥有足够的数据样本是训练高精度DNN作为判别模型的关键点。这种障碍使得DNN不适合数据量有限的许多应用。在这项研究中,我们介绍了一种方法来产生大量的伪样本,只需要少量的转录数据从目标域的可用性。在该方法中,训练通用背景模型(UBM)以捕获数据分布的参数估计。接下来,使用随机采样从UBM生成大量伪样本。然后应用帧混洗来平滑所生成的伪样本序列中的时间倒谱轨迹,以更好地类似于自然语音信号的时间特性。最后,将伪样本序列与原始训练数据相结合来训练语音识别器的DNN-HMM声学模型。所提出的方法进行评估小规模的中性和耳语数据集从UT声乐努力II语料库。结果表明,基于DNN-HMM的语音识别器的音素错误率(PER)大大降低时,将生成的伪样本在训练过程中,与+ 79.0和+ 45.6%的相对PER改善中性中性训练/测试和耳语耳语训练/测试的情况下,分别。
State-of-the-art speech recognition solutions currently employ hidden Markov models (HMMs) to capture the time variability in a speech signal and deep neural networks (DNNs) to model the HMM state distributions. It has been shown that DNN-HMM hybrid systems outperform traditional HMM and Gaussian mixture model (GMM) hybrids in many applications. This improvement is mainly attributed to the ability of DNNs to model more complex data structures. However, having sufficient data samples is one key point in training a high accuracy DNN as a discriminative model. This barrier makes DNNs unsuitable for many applications with limited amounts of data. In this study, we introduce a method to produce an excessive amount of pseudo-samples that requires availability of only a small amount of transcribed data from the target domain. In this method, a universal background model (UBM) is trained to capture a parametric estimate of the data distributions. Next, random sampling is used to generate a large amount of pseudo-samples from the UBM. Frame-Shuffling is then applied to smooth the temporal cepstral trajectories in the generated pseudo-sample sequences to better resemble the temporal characteristics of a natural speech signal. Finally, the pseudo-sample sequences are combined with the original training data to train the DNN-HMM acoustic model of a speech recognizer. The proposed method is evaluated on small-sized sets of neutral and whisper data drawn from the UT-Vocal Effort II corpus. It is shown that phoneme error rates (PERs) of a DNN-HMM based speech recognizer are considerably reduced when incorporating the generated pseudo-samples in the training process, with + 79.0 and + 45.6% relative PER improvements for neutral-neutral training/test and whisper-whisper training/test scenarios, respectively.