Recurrent Neural Network Regularization

Recurrent Neural Network Regularization
复制标题

DOI:
--
复制
发表时间:
2014-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Wojciech Zaremba;I. Sutskever;O. Vinyals
Wojciech Zaremba;I. Sutskever;O. Vinyals
中科院分区:
其他
文献类型:
--
作者:
Wojciech Zaremba;I. Sutskever;O. Vinyals

文献摘要

被引文献

相似文献

我们为具有长短期记忆(LSTM)单元的递归神经网络(RNN)提出了一种简单的正则化技术。Dropout是正则化神经网络的最成功的技术,但在RNN和LSTM中效果不佳。在本文中,我们展示了如何正确地将dropout应用于LSTM,并表明它大大减少了各种任务的过拟合。这些任务包括语言建模、语音识别、图像字幕生成和机器翻译。
We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units. Dropout, the most successful technique for regularizing neural networks, does not work well with RNNs and LSTMs. In this paper, we show how to correctly apply dropout to LSTMs, and show that it substantially reduces overfitting on a variety of tasks. These tasks include language modeling, speech recognition, image caption generation, and machine translation.