Evolution Strategy Based Neural Network Optimization and LSTM Language Model for Robust Speech Recognition

Evolution Strategy Based Neural Network Optimization and LSTM Language Model for Robust Speech Recognition
复制标题

DOI:
--
复制
发表时间:
2016
期刊:
--
影响因子:
--
通讯作者:
Tomohiro Tanaka;T. Shinozaki;Shinji Watanabe;Takaaki Hori
Tomohiro Tanaka;T. Shinozaki;Shinji Watanabe;Takaaki Hori
中科院分区:
其他
文献类型:
--
作者:
Tomohiro Tanaka;T. Shinozaki;Shinji Watanabe;Takaaki Hori

文献摘要

相似文献

本文报道了我们的系统在第四届CHiME挑战赛(CHiME 4)的1通道跟踪任务。开发基于神经网络的系统的一个瓶颈是元参数的调整。我们使用协方差矩阵自适应进化策略(CMA-ES)自动化,使高性能的系统,而不依赖于人类专家。我们对官方基线系统中使用的DNN声学模型进行了两次进化实验。一种是以基于交叉熵(CE)训练后的发展集误词率(WER)作为进化的目标函数,另一种是以顺序判别训练后的WER作为进化的目标函数。此外,我们对基于长短期记忆递归神经网络的语言模型(LSTM-LM)进行了进化实验,取代了基线系统中使用的原始递归神经网络语言模型(RNN-LM),以进行N-最佳重新评分。所有这些进化实验都导致了WER的减少。为了产生最终结果,我们通过汇集来自所有6个通道的语音数据来增强训练数据,并导入优化的元参数设置而无需修改。对于真实的测试数据,当使用RNN和LSTM-LM时,与基线WER 22.75%相比,分别获得了17.40%和16.58%的降低的WER。
This paper reports our system for the 1-channel track task in the 4th CHiME challenge (CHiME4). A bottle-neck in developing neural network based systems is the tuning of meta-parameters. We automate it by using Covariance Matrix Adaptation Evolution Strategy (CMA-ES) so that high performance system is obtained without relying on human experts. We run two evolution experiments for the DNN acoustic model used in the offi-cial baseline system. One uses development set word error rate (WER) after the cross-entropy (CE) based training as the ob-jective function for the evolution, and the other uses the WER after the sequential discriminative training. Additionally, we run an evolution experiment for a Long Short-Term Memory recurrent neural network based language model (LSTM-LM), replacing the original recurrent neural network language model (RNN-LM) used in the baseline system for N-best rescoring. All of these evolution experiments resulted in reduced WERs. To produce the final results, we augmented training data by pooling speech data from all the 6 channels and imported the optimized meta-parameter settings without modification. For the real test data, reduced WER of 17.40% and 16.58% were obtained compared to the baseline WER of 22.75% when the RNN and LSTM-LMs were used, respectively.