Best k-Layer Neural Network Approximations

Best k-Layer Neural Network Approximations
复制标题

最佳 k 层神经网络近似

DOI:
10.1007/s00365-021-09545-2
复制
发表时间:
2022
影响因子:
2.7
通讯作者:
Qi, Yang
Qi, Yang
中科院分区:
数学2区
文献类型:
--
作者:
Lim, Lek-Heng;Michałek, Mateusz;Qi, Yang

文献摘要

参考文献

被引文献

相似文献

我们证明了神经网络的经验风险最小化(ERM)问题一般没有解。给定具有相应响应的训练集,拟合AK层神经网络涉及到通过ERM估计权重:\DocumentClass[12pt]{Minimal}\usepackage{amsath}\usepackage{wa ysym}\usepackage{amsfonts}\usepackage{amsbsy}\usepackage{matrsfs}\usepackage{upgreek}\setlong{\oddsidemarin}{-69pt}\Begin{Document}$\Begin{Begin}\inf_{\theta\in{\mathbb{R}}^m}\\sum_{i=1}^n\Vert t_i-\nu_\theta(S_I)\Vert_2^2。对于RELU、双曲正切和Sigmoid函数等常见的激活,一般不能达到这个下确界。此外,我们还推导出,当损失函数的下确界不能达到时,如果试图最小化该损失函数,必然会导致发散到。我们将展示对于平滑激活,\Docentclass[12pt]{Minimum}\usepackage{amsath}\usepackage{wa ysym}\usepackage{amsfonts}\usepackage{amssymb}\usepackage{amsbsy}\usepackage{mathsfs}\usepackage{upgreek}\setlong{\oddsidemargin}{-69pt}\Begin{Document}$$\sigma(X)=1/\Bigl(1+\exp(-x)\BiGR)$\end{Document}和这种未能达到下确界的情况可能发生在积极测量的反应子集上。对于RELU激活,我们对最佳两层神经网络近似的ERM达到其下确界的情况进行了完全分类。在最近的神经网络应用中,过拟合是很常见的,通过确保方程组有解来避免无法达到下确界的情况。对于两层再激活网络,我们将展示这样的方程组何时有一般解,即这样的神经网络何时可以与概率1完美地拟合。
We show that the empirical risk minimization (ERM) problem for neural networks has no solution in general. Given a training setwith corresponding responses, fitting ak-layer neural networkinvolves estimation of the weightsvia an ERM: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} \inf _{\theta \in {\mathbb {R}}^m} \ \sum _{i=1}^n \Vert t_i - \nu _\theta (s_i) \Vert _2^2. \end{aligned}$$\end{document}We show that even for, this infimum is not attainable in general for common activations like ReLU, hyperbolic tangent, and sigmoid functions. In addition, we deduce that if one attempts to minimize such a loss function in the event when its infimum is not attainable, it necessarily results in values ofdiverging to. We will show that for smooth activations \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma (x)= 1/\bigl (1 + \exp (-x)\bigr )$$\end{document} and, such failure to attain an infimum can happen on a positive-measured subset of responses. For the ReLU activation, we completely classify cases where the ERM for a best two-layer neural network approximation attains its infimum. In recent applications of neural networks, where overfitting is commonplace, the failure to attain an infimum is avoided by ensuring that the system of equations,, has a solution. For a two-layer ReLU-activated network, we will show when such a system of equations has a solution generically, i.e., when can such a neural network be fitted perfectly with probability one.
DOI: 10.1007/978-3-642-38896-5
发表时间: 2013-08
期刊: --
影响因子: --
作者:
Peter Bürgisser;F. Cucker
通讯作者: Peter Bürgisser;F. Cucker
DOI: 10.1090/mmono/127
发表时间: 1993-08
期刊: --
影响因子: --
作者:
F. Zak
通讯作者: F. Zak
张量网络排名
DOI: --
发表时间: 2018
期刊:
影响因子: --
作者:
Ke Ye;Lek
通讯作者: Lek
复杂的最佳 r 项近似几乎总是存在于有限维度中
DOI: 10.1016/j.acha.2018.12.003
发表时间: 2017
影响因子: 2.5
作者:
Yang Qi;M. Michałek;Lek
通讯作者: Lek