Automated structure discovery and parameter tuning of neural network language model based on evolution strategy

Automated structure discovery and parameter tuning of neural network language model based on evolution strategy
复制标题

DOI:
10.1109/slt.2016.7846334
复制
发表时间:
2016-12
期刊:
2016 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Tomohiro Tanaka;Takafumi Moriya;T. Shinozaki;Shinji Watanabe;Takaaki Hori;Kevin Duh
Tomohiro Tanaka;Takafumi Moriya;T. Shinozaki;Shinji Watanabe;Takaaki Hori;Kevin Duh
中科院分区:
其他
文献类型:
--
作者:
Tomohiro Tanaka;Takafumi Moriya;T. Shinozaki;Shinji Watanabe;Takaaki Hori;Kevin Duh

文献摘要

相似文献

众所周知,基于长短期记忆 (LSTM) 循环神经网络的语言模型可以提高语音识别性能。然而,需要付出巨大的努力来优化网络结构和训练配置。在这项研究中,我们使用进化算法自动化开发过程。特别是,我们应用了协方差矩阵适应进化策略(CMA-ES),该策略在其他黑盒超参数优化问题中表现出了鲁棒性。通过灵活地允许优化各种元参数(包括分层单元类型),我们的方法自动找到可提高识别性能的配置。此外,通过使用基于 Pareto 的多目标 CMA-ES,WER 和计算时间都得到了联合减少:在 10 代之后,与 WER 为 8.7% 的初始基线系统相比,解码的相对 WER 和计算时间分别减少了 4.1% 和 22.7%。
Long short-term memory (LSTM) recurrent neural network based language models are known to improve speech recognition performance. However, significant effort is required to optimize network structures and training configurations. In this study, we automate the development process using evolutionary algorithms. In particular, we apply the covariance matrix adaptation-evolution strategy (CMA-ES), which has demonstrated robustness in other black box hyper-parameter optimization problems. By flexibly allowing optimization of various meta-parameters including layer wise unit types, our method automatically finds a configuration that gives improved recognition performance. Further, by using a Pareto based multi-objective CMA-ES, both WER and computational time were reduced jointly: after 10 generations, relative WER and computational time reductions for decoding were 4.1% and 22.7% respectively, compared to an initial baseline system whose WER was 8.7%.