Data Augmentation for Deep Neural Network Acoustic Modeling

Data Augmentation for Deep Neural Network Acoustic Modeling
复制标题

DOI:
10.1109/taslp.2015.2438544
复制
发表时间:
2015-09-01
影响因子:
5.4
通讯作者:
Kingsbury, Brian
Kingsbury, Brian
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cui, Xiaodong;Goel, Vaibhava;Kingsbury, Brian

文献摘要

被引文献

相似文献

本文研究了基于标签保持变换的深度神经网络声学建模的数据增强,以处理数据稀疏性。研究了两种数据增强方法,声道长度扰动(VTLP)和随机特征映射(SFM),用于深度神经网络(DNN)和卷积神经网络(CNN)。这些方法集中于增加有限训练数据的说话者和语音变化,使得用增强数据训练的声学模型对这种变化更鲁棒。此外,一个两阶段的数据增强计划的基础上堆叠架构提出了联合收割机VTLP和SFM作为互补的方法。实验进行了阿萨姆语和海地克里奥尔语,两个开发语言的IARPA巴贝尔计划,并提高性能的自动语音识别(ASR)和关键字搜索(KWS)的报告。
This paper investigates data augmentation for deep neural network acoustic modeling based on label-preserving transformations to deal with data sparsity. Two data augmentation approaches, vocal tract length perturbation (VTLP) and stochastic feature mapping (SFM), are investigated for both deep neural networks (DNNs) and convolutional neural networks (CNNs). The approaches are focused on increasing speaker and speech variations of the limited training data such that the acoustic models trained with the augmented data are more robust to such variations. In addition, a two-stage data augmentation scheme based on a stacked architecture is proposed to combine VTLP and SFM as complementary approaches. Experiments are conducted on Assamese and Haitian Creole, two development languages of the IARPA Babel program, and improved performance on automatic speech recognition (ASR) and keyword search (KWS) is reported.