Feature selection with gradient descent on two-layer networks in low-rotation regimes

Feature selection with gradient descent on two-layer networks in low-rotation regimes
复制标题

DOI:
10.48550/arxiv.2208.02789
复制
发表时间:
2022-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Matus Telgarsky
Matus Telgarsky
中科院分区:
其他
文献类型:
--
作者:
Matus Telgarsky

文献摘要

被引文献

相似文献

这项工作在具有标准初始化的两层 ReLU 网络上建立了梯度流 (GF) 和随机梯度下降 (SGD) 的低测试误差,在关键权重集旋转很少的三种状态下(自然地由于 GF 和 SGD,或者由于人为约束),并利用边际作为核心分析技术。第一个状态接近初始化,特别是直到权重移动 $\mathcal{O}(\sqrt m)$,其中 $m$ 表示网络宽度,这与神经正切核 (NTK) 允许的 $\mathcal{O}(1)$ 权重运动形成鲜明对比;这里表明,GF和SGD只需要与NTK裕度成反比的网络宽度和样本数量,而且GF至少达到NTK裕度本身,这足以建立逃离裕度目标的不良KKT点的方法,而先前的工作只能建立非递减但任意小的裕度。第二种状态是神经崩溃(NC)设置,其中数据位于分离得非常好的组中,并且样本复杂性随着组的数量而变化;这里对先前工作的贡献是对初始化时整个 GF 轨迹的分析。最后,如果内层权重被限制为仅在范数上变化并且不能旋转,则具有大宽度的GF实现全局最大边际,并且其样本复杂度与其倒数成比例;这与之前的工作形成鲜明对比,之前的工作需要无限的宽度和棘手的对偶收敛假设。作为纯粹的技术贡献,这项工作开发了各种潜在的功能和其他工具,希望对未来的工作有所帮助。
This work establishes low test error of gradient flow (GF) and stochastic gradient descent (SGD) on two-layer ReLU networks with standard initialization, in three regimes where key sets of weights rotate little (either naturally due to GF and SGD, or due to an artificial constraint), and making use of margins as the core analytic technique. The first regime is near initialization, specifically until the weights have moved by $\mathcal{O}(\sqrt m)$, where $m$ denotes the network width, which is in sharp contrast to the $\mathcal{O}(1)$ weight motion allowed by the Neural Tangent Kernel (NTK); here it is shown that GF and SGD only need a network width and number of samples inversely proportional to the NTK margin, and moreover that GF attains at least the NTK margin itself, which suffices to establish escape from bad KKT points of the margin objective, whereas prior work could only establish nondecreasing but arbitrarily small margins. The second regime is the Neural Collapse (NC) setting, where data lies in extremely-well-separated groups, and the sample complexity scales with the number of groups; here the contribution over prior work is an analysis of the entire GF trajectory from initialization. Lastly, if the inner layer weights are constrained to change in norm only and can not rotate, then GF with large widths achieves globally maximal margins, and its sample complexity scales with their inverse; this is in contrast to prior work, which required infinite width and a tricky dual convergence assumption. As purely technical contributions, this work develops a variety of potential functions and other tools which will hopefully aid future work.