A Theoretically Grounded Application of Dropout in Recurrent Neural Networks

A Theoretically Grounded Application of Dropout in Recurrent Neural Networks
复制标题

DOI:
--
复制
发表时间:
2015-12
期刊:
--
影响因子:
--
通讯作者:
Y. Gal;Zoubin Ghahramani
Y. Gal;Zoubin Ghahramani
中科院分区:
其他
文献类型:
--
作者:
Y. Gal;Zoubin Ghahramani

文献摘要

被引文献

相似文献

递归神经网络(RNN)处于深度学习许多最新发展的最前沿。然而,这些模型的一个主要困难是它们倾向于过拟合,当应用于递归层时,dropout会失败。贝叶斯建模和深度学习交叉的最新结果为常见的深度学习技术(如dropout)提供了贝叶斯解释。这种近似贝叶斯推理中的dropout基础表明了理论结果的扩展,为RNN模型中dropout的使用提供了见解。我们在LSTM和GRU模型中应用了这种新的基于变分推理的dropout技术,并在语言建模和情感分析任务中对其进行了评估。新的方法优于现有的技术,并尽我们所知,提高了单一模型的最先进的语言建模与Penn Treebank(73.4测试困惑)。这扩展了我们在深度学习中的变分工具库。
Recurrent neural networks (RNNs) stand at the forefront of many recent developments in deep learning. Yet a major difficulty with these models is their tendency to overfit, with dropout shown to fail when applied to recurrent layers. Recent results at the intersection of Bayesian modelling and deep learning offer a Bayesian interpretation of common deep learning techniques such as dropout. This grounding of dropout in approximate Bayesian inference suggests an extension of the theoretical results, offering insights into the use of dropout with RNN models. We apply this new variational inference based dropout technique in LSTM and GRU models, assessing it on language modelling and sentiment analysis tasks. The new approach outperforms existing techniques, and to the best of our knowledge improves on the single model state-of-the-art in language modelling with the Penn Treebank (73.4 test perplexity). This extends our arsenal of variational tools in deep learning.