Best k-Layer Neural Network Approximations
Best k-Layer Neural Network Approximations
复制标题
最佳 k 层神经网络近似
DOI:
10.1007/s00365-021-09545-2
复制
发表时间:
2022
影响因子:
2.7
通讯作者:
Qi, Yang
中科院分区:
文献类型:
--
作者:
Lim, Lek-Heng;Michałek, Mateusz;Qi, Yang
We show that the empirical risk minimization (ERM) problem for neural networks has no solution in general. Given a training setwith corresponding responses, fitting ak-layer neural networkinvolves estimation of the weightsvia an ERM: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\begin{aligned} \inf _{\theta \in {\mathbb {R}}^m} \ \sum _{i=1}^n \Vert t_i - \nu _\theta (s_i) \Vert _2^2. \end{aligned}$$\end{document}We show that even for, this infimum is not attainable in general for common activations like ReLU, hyperbolic tangent, and sigmoid functions. In addition, we deduce that if one attempts to minimize such a loss function in the event when its infimum is not attainable, it necessarily results in values ofdiverging to. We will show that for smooth activations \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\sigma (x)= 1/\bigl (1 + \exp (-x)\bigr )$$\end{document} and, such failure to attain an infimum can happen on a positive-measured subset of responses. For the ReLU activation, we completely classify cases where the ERM for a best two-layer neural network approximation attains its infimum. In recent applications of neural networks, where overfitting is commonplace, the failure to attain an infimum is avoided by ensuring that the system of equations,, has a solution. For a two-layer ReLU-activated network, we will show when such a system of equations has a solution generically, i.e., when can such a neural network be fitted perfectly with probability one.
登录
查看更多内容
DOI:
10.1007/978-3-642-38896-5
发表时间:
2013-08
期刊:
--
影响因子:
--
作者:
Peter Bürgisser;F. Cucker
通讯作者:
Peter Bürgisser;F. Cucker
DOI:
10.1090/mmono/127
发表时间:
1993-08
期刊:
--
影响因子:
--
作者:
F. Zak
通讯作者:
F. Zak
DOI:
--
发表时间:
2018
期刊:
影响因子:
--
作者:
Ke Ye;Lek
通讯作者:
Lek
影响因子:
2.5
作者:
Yang Qi;M. Michałek;Lek
通讯作者:
Lek