Neural network approximation

Neural network approximation
复制标题

DOI:
10.1017/s0962492921000052
复制
发表时间:
2021-05-01
期刊:
影响因子:
14.2
通讯作者:
Petrova, Guergana
Petrova, Guergana
中科院分区:
数学1区
文献类型:
--
作者:
DeVore, Ronald;Hanin, Boris;Petrova, Guergana

文献摘要

被引文献

相似文献

神经网络(NN)是构建学习算法的首选方法。他们现在正在研究其他数值任务,如解决高维偏微分方程。他们的受欢迎程度源于他们在几个具有挑战性的学习问题(计算机象棋/围棋,自主导航,人脸识别)上的经验成功。然而,大多数学者认为,这一成功仍然缺乏令人信服的理论解释。由于这些应用都是围绕着从数据观测中近似未知函数,因此部分答案必须涉及NN产生精确近似的能力。本文调查了NN输出的已知逼近性质,旨在揭示数值分析中使用的更传统逼近方法中不存在的性质,例如使用多项式、小波、有理函数和样条的逼近。从速率失真的角度,即误差与用于创建近似的参数的数量与传统的近似方法进行比较。在分析数值逼近的另一个主要组成部分是计算时间需要构建的近似,这反过来又是密切相关的近似算法的稳定性。因此,使用神经网络的数值逼近的稳定性是提出的分析的很大一部分。该调查在很大程度上与使用流行的ReLU激活功能的NN有关。在这种情况下,神经网络的输出是分段线性函数,它将f的域划分为凸多面体的单元。当神经网络的结构是固定的,并允许参数变化时,神经网络的输出函数集是一个参数化的非线性流形。结果表明,该流形具有一定的空间填充特性,从而提高了近似的能力(更好的率失真),但以牺牲数值稳定性为代价。当试图近似时,空间填充对找到最佳或良好参数选择的数值方法产生了挑战。
Neural networks (NNs) are the method of choice for building learning algorithms. They are now being investigated for other numerical tasks such as solving high-dimensional partial differential equations. Their popularity stems from their empirical success on several challenging learning problems (computer chess/Go, autonomous navigation, face recognition). However, most scholars agree that a convincing theoretical explanation for this success is still lacking. Since these applications revolve around approximating an unknown function from data observations, part of the answer must involve the ability of NNs to produce accurate approximations. This article surveys the known approximation properties of the outputs of NNs with the aim of uncovering the properties that are not present in the more traditional methods of approximation used in numerical analysis, such as approximations using polynomials, wavelets, rational functions and splines. Comparisons are made with traditional approximation methods from the viewpoint of rate distortion, i.e. error versus the number of parameters used to create the approximant. Another major component in the analysis of numerical approximation is the computational time needed to construct the approximation, and this in turn is intimately connected with the stability of the approximation algorithm. So the stability of numerical approximation using NNs is a large part of the analysis put forward. The survey, for the most part, is concerned with NNs using the popular ReLU activation function. In this case the outputs of the NNs are piecewise linear functions on rather complicated partitions of the domain of f into cells that are convex polytopes. When the architecture of the NN is fixed and the parameters are allowed to vary, the set of output functions of the NN is a parametrized nonlinear manifold. It is shown that this manifold has certain space-filling properties leading to an increased ability to approximate (better rate distortion) but at the expense of numerical stability. The space filling creates the challenge to the numerical method of finding best or good parameter choices when trying to approximate.