Data Augmentation for deep neural network acoustic modeling

Data Augmentation for deep neural network acoustic modeling
复制标题

DOI:
10.1109/icassp.2014.6854671
复制
发表时间:
2014-05
期刊:
2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Xiaodong Cui;Vaibhava Goel;Brian Kingsbury
Xiaodong Cui;Vaibhava Goel;Brian Kingsbury
中科院分区:
其他
文献类型:
--
作者:
Xiaodong Cui;Vaibhava Goel;Brian Kingsbury

文献摘要

被引文献

相似文献

使用标签保持变换的数据增强已被证明对神经网络训练进行不变预测是有效的。在本文中,我们专注于使用深度神经网络(DNN)进行自动语音识别(ASR)的声学建模的数据增强方法。我们首先研究了一个修改版本的先前研究的方法,使用声道长度扰动(VTLP),然后提出了一种新的数据增强方法的基础上随机特征映射(SFM)在扬声器自适应特征空间。实验进行孟加拉语和阿萨姆语有限的语言包(LLP)从IARPA巴贝尔计划。在DNN模型的交叉熵(CE)和状态级最小贝叶斯风险(sMBR)训练之后,已经观察到识别性能的改善。
Data augmentation using label preserving transformations has been shown to be effective for neural network training to make invariant predictions. In this paper we focus on data augmentation approaches to acoustic modeling using deep neural networks (DNNs) for automatic speech recognition (ASR). We first investigate a modified version of a previously studied approach using vocal tract length perturbation (VTLP) and then propose a novel data augmentation approach based on stochastic feature mapping (SFM) in a speaker adaptive feature space. Experiments were conducted on Bengali and Assamese limited language packs (LLPs) from the IARPA Babel program. Improved recognition performance has been observed after both cross-entropy (CE) and state-level minimum Bayes risk (sMBR) training of DNN models.