When Expressivity Meets Trainability: Fewer than $n$ Neurons Can Work

When Expressivity Meets Trainability: Fewer than $n$ Neurons Can Work
复制标题

DOI:
10.48550/arxiv.2210.12001
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jiawei Zhang;Yushun Zhang;Mingyi Hong;Ruoyu Sun;Z. Luo
Jiawei Zhang;Yushun Zhang;Mingyi Hong;Ruoyu Sun;Z. Luo
中科院分区:
其他
文献类型:
--
作者:
Jiawei Zhang;Yushun Zhang;Mingyi Hong;Ruoyu Sun;Z. Luo

文献摘要

相似文献

现代神经网络通常是相当广泛的,导致大量的内存和计算成本。因此,训练一个更窄的网络是非常有趣的。然而,训练窄神经网络仍然是一项具有挑战性的任务。我们提出了两个理论上的问题:狭义网络能像广义网络一样具有强大的表现力吗?如果是这样,损失函数是否呈现出良性的优化景观?在这项工作中,我们为激活平滑时神经元少于$n$(样本量)的1隐藏层网络提供了部分肯定的答案。首先,我们证明了只要宽度为$m \geq 2n/d$(其中$d$为输入维数),其表达性就很强,即存在至少一个训练损失为零的全局最小化器。其次,我们确定一个很好的局部区域,没有local-min或鞍点。然而,梯度下降是否能在这一良好区域内持续存在,目前尚不清楚。第三,我们考虑了一个可行区域为优局部区域的约束优化公式,并证明了每一个KKT点都是一个近全局最小化点。预计投影梯度方法在温和的技术条件下收敛到KKT点,但我们将严格的收敛分析留给未来的工作。全面的数值结果表明,在此约束公式上的投影梯度方法在训练窄神经网络方面明显优于SGD方法。
Modern neural networks are often quite wide, causing large memory and computation costs. It is thus of great interest to train a narrower network. However, training narrow neural nets remains a challenging task. We ask two theoretical questions: Can narrow networks have as strong expressivity as wide ones? If so, does the loss function exhibit a benign optimization landscape? In this work, we provide partially affirmative answers to both questions for 1-hidden-layer networks with fewer than $n$ (sample size) neurons when the activation is smooth. First, we prove that as long as the width $m \geq 2n/d$ (where $d$ is the input dimension), its expressivity is strong, i.e., there exists at least one global minimizer with zero training loss. Second, we identify a nice local region with no local-min or saddle points. Nevertheless, it is not clear whether gradient descent can stay in this nice region. Third, we consider a constrained optimization formulation where the feasible region is the nice local region, and prove that every KKT point is a nearly global minimizer. It is expected that projected gradient methods converge to KKT points under mild technical conditions, but we leave the rigorous convergence analysis to future work. Thorough numerical results show that projected gradient methods on this constrained formulation significantly outperform SGD for training narrow neural nets.