On Generalization Bounds of a Family of Recurrent Neural Networks

On Generalization Bounds of a Family of Recurrent Neural Networks
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Minshuo Chen;Xingguo Li;T. Zhao
Minshuo Chen;Xingguo Li;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Minshuo Chen;Xingguo Li;T. Zhao

文献摘要

相似文献

循环神经网络(RNN)已广泛应用于顺序数据分析。然而,由于其复杂的建模结构,背后的理论仍然很大程度上缺失。为了将理论与实践联系起来,我们研究了普通 RNN 及其变体的泛化特性,包括最小门控单元 (MGU)、长短期记忆 (LSTM) 和卷积 (Conv) RNN。具体来说,我们的理论是在 PAC-Learning 框架下建立的。泛化界限以权重矩阵的谱范数和参数总数的形式表示。我们还通过附加规范假设建立了精炼的泛化界限,并对这些界限进行了比较。我们评论道:(1)我们对普通 RNN 的泛化界限明显比现有最好的结果更严格; (2) 我们不知道现有文献中 MGU、LSTM 和 Conv RNN 的任何其他泛化界限; (3)我们证明了这些变体在泛化方面的优势。
Recurrent Neural Networks (RNNs) have been widely applied to sequential data analysis. Due to their complicated modeling structures, however, the theory behind is still largely missing. To connect theory and practice, we study the generalization properties of vanilla RNNs as well as their variants, including Minimal Gated Unit (MGU), Long Short Term Memory (LSTM), and Convolutional (Conv) RNNs. Specifically, our theory is established under the PAC-Learning framework. The generalization bound is presented in terms of the spectral norms of the weight matrices and the total number of parameters. We also establish refined generalization bounds with additional norm assumptions, and draw a comparison among these bounds. We remark: (1) Our generalization bound for vanilla RNNs is significantly tighter than the best of existing results; (2) We are not aware of any other generalization bounds for MGU, LSTM, and Conv RNNs in the exiting literature; (3) We demonstrate the advantages of these variants in generalization.