Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods

Excess Risk of Two-Layer ReLU Neural Networks in Teacher-Student Settings and its Superiority to Kernel Methods
复制标题

DOI:
10.48550/arxiv.2205.14818
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Shunta Akiyama;Taiji Suzuki
Shunta Akiyama;Taiji Suzuki
中科院分区:
其他
文献类型:
--
作者:
Shunta Akiyama;Taiji Suzuki

文献摘要

相似文献

虽然深度学习在各种任务中的表现优于其他方法,但解释其原因的理论框架尚未完全建立。为了解决这个问题,我们研究了两层RELU神经网络在师生回归模型中的额外风险,在该模型中,学生网络通过其输出学习未知的教师网络。特别是,我们考虑了与教师网络具有相同宽度的学生网络,并分两个阶段进行训练:首先是噪声梯度下降,然后是香草梯度下降。我们的结果表明,在极小极大最优率的意义下,学生网络可以证明达到近全局最优解,并且性能优于任何核方法估计器(更一般地,线性估计器),包括神经切核方法、随机特征模型和其他核方法。导致这种优势的关键概念是神经网络模型的非凸性。即使损失情况是高度非凸的,学生网络也会自适应地学习老师的神经元。
While deep learning has outperformed other methods for various tasks, theoretical frameworks that explain its reason have not been fully established. To address this issue, we investigate the excess risk of two-layer ReLU neural networks in a teacher-student regression model, in which a student network learns an unknown teacher network through its outputs. Especially, we consider the student network that has the same width as the teacher network and is trained in two phases: first by noisy gradient descent and then by the vanilla gradient descent. Our result shows that the student network provably reaches a near-global optimal solution and outperforms any kernel methods estimator (more generally, linear estimators), including neural tangent kernel approach, random feature model, and other kernel methods, in a sense of the minimax optimal rate. The key concept inducing this superiority is the non-convexity of the neural network models. Even though the loss landscape is highly non-convex, the student network adaptively learns the teacher neurons.