Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis

Benefits of Jointly Training Autoencoders: An Improved Neural Tangent Kernel Analysis
复制标题

DOI:
10.1109/tit.2021.3065212
复制
发表时间:
2019-11
影响因子:
2.5
通讯作者:
THANH VAN NGUYEN;Raymond K. W. Wong;C. Hegde
THANH VAN NGUYEN;Raymond K. W. Wong;C. Hegde
中科院分区:
计算机科学2区
文献类型:
--
作者:
THANH VAN NGUYEN;Raymond K. W. Wong;C. Hegde

文献摘要

被引文献

相似文献

在深度神经网络被大量过度参数化的情况下,深度神经网络可以取得令人印象深刻的性能。因此,在过去的一年里,人们对分析过参数网络的最优化和泛化性质越来越感兴趣。然而,现有的大部分工作只适用于有监督的学习。相比之下,过度参数化在无监督环境中的作用得到的关注要少得多。本文研究了具有RELU激活的双层超参数自动编码器的梯度下降的感应偏差。我们首先利用无限宽的神经网络在梯度下降下演化为线性模型的性质,为最近的工作中观察到的记忆现象提供了理论证据。我们还分析了有限宽度环境下自动编码器的梯度动态特性。从一个随机初始化的自动编码网络出发,严格证明了梯度下降在弱训练和联合训练两种情况下的线性收敛。我们的结果表明,联合训练比弱训练在寻找全局最优解方面有相当大的好处,实现了所需的过度参数化水平的显著降低。最后,我们分析了权重绑定自动编码器的情况,证明了在过参数设置下,从随机初始化点训练这种网络会导致某些意想不到的退化。
Deep neural networks can achieve impressive performance in the regime where they are massively over-parameterized. Consequently, over the past year, there has been a growing interest in analyzing optimization and generalization properties of over-parameterized networks. However, the majority of existing work only applies to supervised learning. The role of over-parameterization in the unsupervised setting has by contrast gained far less attention. In this paper, we study the inductive bias of gradient descent for two-layer over-parameterized autoencoders with ReLU activation. We first provide theoretical evidence for the memorization phenomena observed in recent work using the property that infinitely wide neural networks under gradient descent evolve as linear models. We also analyze the gradient dynamics of the autoencoders in the finite-width setting. Starting from a randomly initialized autoencoder network, we rigorously prove the linear convergence of gradient descent in two weakly-trained and jointly-trained regimes. Our results indicate the considerable benefits of joint training over weak training in finding global optima, achieving a dramatic decrease in the required level of over-parameterization. Finally, we analyze the case of weight-tied autoencoders and prove that in the over-parameterized setting, training such networks from randomly initialized points leads to certain unexpected degeneracies.