Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits

Training Neural Networks as Learning Data-adaptive Kernels: Provable Representation and Approximation Benefits
复制标题

DOI:
10.1080/01621459.2020.1745812
复制
发表时间:
2019-01
影响因子:
3.7
通讯作者:
Xialiang Dou;Tengyuan Liang
Xialiang Dou;Tengyuan Liang
中科院分区:
数学1区
文献类型:
--
作者:
Xialiang Dou;Tengyuan Liang

文献摘要

被引文献

相似文献

摘要 考虑这个问题:给定从 总体中提取的数据对,指定一个神经网络模型,并随着时间的推移对权重运行梯度流,直到达到平稳性。 ft (由神经网络在时间 t 计算的函数)在近似和表示方面与 有何关系?与经典非参数文献中预先指定的固定基础表示相比,神经网络的自适应表示有哪些可证明的好处?我们通过神经网络训练过程索引的动态再现核希尔伯特空间(RKHS)方法回答了上述问题。首先,我们表明,当达到任何局部平稳性时,梯度流学习自适应 RKHS 表示,并同时在自适应 RKHS 上执行全局最小二乘投影。其次,我们证明由于 RKHS 是数据自适应且特定于任务的,因此 的残差位于可能比 RKHS 的正交补小得多的子空间中。结果形式化了神经网络的表示和近似优势。最后,我们证明了在消失正则化的极限下,由梯度流计算的神经网络函数收敛到具有自适应核的核无脊回归。自适应核的观点为研究神经网络的逼近、表示、泛化和优化优势提供了新的角度。
Abstract Consider the problem: given the data pair drawn from a population with , specify a neural network model and run gradient flow on the weights over time until reaching any stationarity. How does ft , the function computed by the neural network at time t, relate to , in terms of approximation and representation? What are the provable benefits of the adaptive representation by neural networks compared to the pre-specified fixed basis representation in the classical nonparametric literature? We answer the above questions via a dynamic reproducing kernel Hilbert space (RKHS) approach indexed by the training process of neural networks. Firstly, we show that when reaching any local stationarity, gradient flow learns an adaptive RKHS representation and performs the global least-squares projection onto the adaptive RKHS, simultaneously. Secondly, we prove that as the RKHS is data-adaptive and task-specific, the residual for lies in a subspace that is potentially much smaller than the orthogonal complement of the RKHS. The result formalizes the representation and approximation benefits of neural networks. Lastly, we show that the neural network function computed by gradient flow converges to the kernel ridgeless regression with an adaptive kernel, in the limit of vanishing regularization. The adaptive kernel viewpoint provides new angles of studying the approximation, representation, generalization, and optimization advantages of neural networks.