Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?

Deep Learning meets Nonparametric Regression: Are Weight-Decayed DNNs Locally Adaptive?
复制标题

DOI:
10.48550/arxiv.2204.09664
复制
发表时间:
2022-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Kaiqi Zhang;Yu-Xiang Wang
Kaiqi Zhang;Yu-Xiang Wang
中科院分区:
其他
文献类型:
--
作者:
Kaiqi Zhang;Yu-Xiang Wang

文献摘要

被引文献

相似文献

我们研究了神经网络(NN)的理论,从经典的非参数回归问题的透镜,重点是NN的能力,自适应估计函数的非均匀光滑性-一个属性的功能在Besov或有界变差(BV)类。现有的工作在这个问题上需要调整的神经网络结构的函数空间和样本大小的基础上。我们考虑深度ReLU网络的“并行NN“变体,并表明标准$\ell_2$正则化相当于提升端到端学习函数基的系数向量中的$\ell_p$-稀疏性($0<p<1$),即,字典使用这种等价性,我们进一步建立,通过调整正则化因子,这样的并行NN实现了任意接近Besov和BV类的极大极小率的估计误差。值得注意的是,随着NN的深入,它以指数方式接近最小最大最优。我们的研究揭示了为什么深度很重要以及NN如何比内核方法更强大。
We study the theory of neural network (NN) from the lens of classical nonparametric regression problems with a focus on NN's ability to adaptively estimate functions with heterogeneous smoothness -- a property of functions in Besov or Bounded Variation (BV) classes. Existing work on this problem requires tuning the NN architecture based on the function spaces and sample size. We consider a"Parallel NN"variant of deep ReLU networks and show that the standard $\ell_2$ regularization is equivalent to promoting the $\ell_p$-sparsity ($0<p<1$) in the coefficient vector of an end-to-end learned function bases, i.e., a dictionary. Using this equivalence, we further establish that by tuning only the regularization factor, such parallel NN achieves an estimation error arbitrarily close to the minimax rates for both the Besov and BV classes. Notably, it gets exponentially closer to minimax optimal as the NN gets deeper. Our research sheds new lights on why depth matters and how NNs are more powerful than kernel methods.