Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data

Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional Data
复制标题

DOI:
10.48550/arxiv.2210.07082
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Spencer Frei;Gal Vardi;P. Bartlett;N. Srebro;Wei Hu
Spencer Frei;Gal Vardi;P. Bartlett;N. Srebro;Wei Hu
中科院分区:
其他
文献类型:
--
作者:
Spencer Frei;Gal Vardi;P. Bartlett;N. Srebro;Wei Hu

文献摘要

被引文献

相似文献

基于梯度的优化算法的隐式偏差被认为是现代深度学习成功的主要因素。在这项工作中,我们研究了当训练数据接近正交时,具有泄漏ReLU激活的两层全连接神经网络中梯度流和梯度下降的隐式偏差,这是高维数据的一个常见属性。对于梯度流,我们利用最近关于同质神经网络隐式偏差的工作来证明,梯度流产生的神经网络的秩最多为2。此外,该网络是一个最大边际解(在参数空间中),并且具有对应于近似最大边际线性预测器的线性决策边界。对于梯度下降,假设随机初始化方差足够小,我们证明了梯度下降的一个步骤足以大幅降低网络的秩,并且在整个训练过程中秩保持很小。我们提供的实验表明,一个小的初始化规模是很重要的寻找低秩神经网络梯度下降。
The implicit biases of gradient-based optimization algorithms are conjectured to be a major factor in the success of modern deep learning. In this work, we investigate the implicit bias of gradient flow and gradient descent in two-layer fully-connected neural networks with leaky ReLU activations when the training data are nearly-orthogonal, a common property of high-dimensional data. For gradient flow, we leverage recent work on the implicit bias for homogeneous neural networks to show that asymptotically, gradient flow produces a neural network with rank at most two. Moreover, this network is an $\ell_2$-max-margin solution (in parameter space), and has a linear decision boundary that corresponds to an approximate-max-margin linear predictor. For gradient descent, provided the random initialization variance is small enough, we show that a single step of gradient descent suffices to drastically reduce the rank of the network, and that the rank remains small throughout training. We provide experiments which suggest that a small initialization scale is important for finding low-rank neural networks with gradient descent.