Cluster adaptive training for deep neural network

Cluster adaptive training for deep neural network
复制标题

DOI:
10.1109/icassp.2015.7178787
复制
发表时间:
2015-04
期刊:
2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu
中科院分区:
其他
文献类型:
--
作者:
Tian Tan;Y. Qian;Maofan Yin;Yimeng Zhuang;Kai Yu

文献摘要

被引文献

相似文献

尽管上下文相关的 DNN-HMM 系统比 GMM-HMM 系统取得了显着的改进,但如果测试数据的声学条件与训练数据的声学条件不匹配,仍然存在很大的性能下降。因此,DNN 的适应和自适应训练具有很大的研究兴趣。之前的工作主要集中在通过正则化或选择性微调来调整单个 DNN 的参数,将线性变换应用于特征或隐藏层输出,或者将非语音可变性的向量表示引入输入中。这些方法都需要在适应过程中估计相对大量的参数。相比之下,本文采用集群自适应训练(CAT)框架进行 DNN 自适应。这里,构建多个 DNN 以形成规范参数空间的基础。在适应过程中,特定于特定声学条件的插值向量用于将多个 DNN 基组合成单个适应的 DNN。 DNN 基础也可以在层级别构建,以获得更大的灵活性。 CAT-DNN 方法在无监督适应模式下的英语总机任务上进行了评估。与未适应的 DNN-HMM 相比,它仅用 10 个参数就显着降低了 WER,相对降低了 6% 到 8.5%。
Although context-dependent DNN-HMM systems have achieved significant improvements over GMM-HMM systems, there still exists big performance degradation if the acoustic condition of the test data mismatches that of the training data. Hence, adaptation and adaptive training of DNN are of great research interest. Previous works mainly focus on adapting the parameters of a single DNN by regularized or selective fine-tuning, applying linear transforms to feature or hidden-layer output, or introducing vector representation of non-speech variability into the input. These methods all require relatively large number of parameters to be estimated during adaptation. In contrast, this paper employs the cluster adaptive training (CAT) framework for DNN adaptation. Here, multiple DNNs are constructed to form the bases of a canonical parametric space. During adaptation, an interpolation vector, specific to a particular acoustic condition, is used to combine the multiple DNN bases into a single adapted DNN. The DNN bases can also be constructed at layer level for more flexibility. The CAT-DNN approach was evaluated on an English switchboard task in unsupervised adaptation mode. It achieved significant WER reductions over the unadapted DNN-HMM, relative 6% to 8.5%, with only 10 parameters.