A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case

A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate Case
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Greg Ongie;R. Willett;Daniel Soudry;N. Srebro
Greg Ongie;R. Willett;Daniel Soudry;N. Srebro
中科院分区:
其他
文献类型:
--
作者:
Greg Ongie;R. Willett;Daniel Soudry;N. Srebro

文献摘要

相似文献

理解过参数化神经网络有效性的一个关键要素是描述当网络中的权重数量接近无穷大时,它们如何表示函数。在本文中,我们将实现函数$f:\mathbb{R}^d\rightarrow\mathbb{R}$所需的范数表征为具有无界单元数(无限宽度)的单个隐藏层ReLU网络,但其中权值的欧几里得范数是有界的,包括精确表征哪些函数可以用有限范数实现。这在Savarese等人(2019)的单变量单变量函数中得到了解决,其中表明所需的范数由函数二阶导数的l1范数决定。我们将表征扩展到多元函数(即,具有d个输入单元的网络),将所需范数与函数的a (d+1)/2次方拉普拉斯的Radon变换的l1范数联系起来。这种表征使我们能够证明Sobolev空间$W^{s,1}(\mathbb{R})$, $s\geq d+1$中的所有函数都可以用有界范数表示,从而计算几个特定函数所需的范数,并获得深度分离结果。这些结果对于理解泛化性能以及神经网络与更传统的核学习之间的区别具有重要意义。
A key element of understanding the efficacy of overparameterized neural networks is characterizing how they represent functions as the number of weights in the network approaches infinity. In this paper, we characterize the norm required to realize a function $f:\mathbb{R}^d\rightarrow\mathbb{R}$ as a single hidden-layer ReLU network with an unbounded number of units (infinite width), but where the Euclidean norm of the weights is bounded, including precisely characterizing which functions can be realized with finite norm. This was settled for univariate univariate functions in Savarese et al. (2019), where it was shown that the required norm is determined by the L1-norm of the second derivative of the function. We extend the characterization to multivariate functions (i.e., networks with d input units), relating the required norm to the L1-norm of the Radon transform of a (d+1)/2-power Laplacian of the function. This characterization allows us to show that all functions in Sobolev spaces $W^{s,1}(\mathbb{R})$, $s\geq d+1$, can be represented with bounded norm, to calculate the required norm for several specific functions, and to obtain a depth separation result. These results have important implications for understanding generalization performance and the distinction between neural networks and more traditional kernel learning.