Banach Space Representer Theorems for Neural Networks and Ridge Splines

Banach Space Representer Theorems for Neural Networks and Ridge Splines
复制标题

DOI:
--
复制
发表时间:
2020-11
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Rahul Parhi;R. Nowak
Rahul Parhi;R. Nowak
中科院分区:
其他
文献类型:
--
作者:
Rahul Parhi;R. Nowak

文献摘要

被引文献

相似文献

我们开发了一个变分框架来理解神经网络学习到的函数的属性适合数据。我们提出并研究了一系列连续域线性逆问题,在受数据拟合约束的 Radon 域中具有类总变分正则化。我们推导出一个表示定理,表明有限宽度、单隐藏层神经网络是这些逆问题的解决方案。我们借鉴了变分样条理论中的许多技术,因此提出了多项式岭样条的概念,它对应于以截断幂函数作为激活函数的单隐藏层神经网络。表示定理让人想起经典的再现核希尔伯特空间表示定理,但我们表明神经网络问题是在非希尔伯特巴纳赫空间上提出的。虽然学习问题是在连续域中提出的,但与核方法类似,这些问题可以重新转换为有限维神经网络训练问题。这些神经网络训练问题具有与众所周知的权重衰减和路径范数正则化器相关的正则化器。因此,我们的结果可以深入了解经过训练的神经网络的功能特征以及设计神经网络正则化器。我们还表明,这些正则化器可以促进具有理想泛化特性的神经网络解决方案。
We develop a variational framework to understand the properties of the functions learned by neural networks fit to data. We propose and study a family of continuous-domain linear inverse problems with total variation-like regularization in the Radon domain subject to data fitting constraints. We derive a representer theorem showing that finite-width, single-hidden layer neural networks are solutions to these inverse problems. We draw on many techniques from variational spline theory and so we propose the notion of polynomial ridge splines, which correspond to a single-hidden layer neural networks with truncated power functions as the activation function. The representer theorem is reminiscent of the classical reproducing kernel Hilbert space representer theorem, but we show that the neural network problem is posed over a non-Hilbertian Banach space. While the learning problems are posed in the continuous-domain, similar to kernel methods, the problems can be recast as finite-dimensional neural network training problems. These neural network training problems have regularizers which are related to the well-known weight decay and path-norm regularizers. Thus, our result gives insight into functional characteristics of trained neural networks and also into the design neural network regularizers. We also show that these regularizers promote neural network solutions with desirable generalization properties.