A unified and constructive framework for the universality of neural networks

A unified and constructive framework for the universality of neural networks
复制标题

神经网络通用性的统一和建设性框架

DOI:
10.1093/imamat/hxad032
复制
发表时间:
2023
影响因子:
1.2
通讯作者:
Bui-Thanh, Tan
Bui-Thanh, Tan
中科院分区:
数学4区
文献类型:
--
作者:
Bui-Thanh, Tan

文献摘要

相似文献

许多神经网络能够复制复杂任务或函数的原因之一是它们的通用逼近性质。尽管过去几十年神经网络理论取得了巨大进步,但神经网络通用性的单一建设性基本框架仍然不存在。本文致力于为一大类激活函数(包括大多数现有激活函数)的通用性提供一个统一且有建设性的框架。该框架的核心是神经网络近似同一性(nAI)的概念。主要结果如下:任何nAI激活函数在compacta上的连续函数空间中都是通用的。事实证明,大多数现有的激活函数都是nAI,因此具有通用性。与当代同类框架相比,该框架具有多个优势。首先,它具有泛函分析、概率论和数值分析的基本手段。其次,它是第一个统一的、建设性的尝试,对大多数现有的激活函数都有效。第三,它为大多数激活函数提供了新的证明。第四,对于给定的激活和容错,该框架精确地提供了具有预定数量的神经元和权重/偏差值的相应单隐神经网络的架构。第五,该框架允许我们抽象地提出具有有利的非渐近率的第一个通用近似。第六,我们的框架还提供了对某些现有方法的发展的见解,从而提供了建设性的推导。
One of the reasons why many neural networks are capable of replicating complicated tasks or functions is their universal approximation property. Though the past few decades have seen tremendous advances in theories of neural networks, a single constructive and elementary framework for neural network universality remains unavailable. This paper is an effort to provide a unified and constructive framework for the universality of a large class of activation functions including most of the existing ones. At the heart of the framework is the concept of neural network approximate identity (nAI). The main result is as follows:any nAI activation function is universal in the space of continuous functions on compacta. It turns out that most of the existing activation functions are nAI, and thus universal. The framework inducesseveral advantagesover the contemporary counterparts. First, it is constructive with elementary means from functional analysis, probability theory, and numerical analysis. Second, it is one of the first unified and constructive attempts that is valid for most of the existing activation functions. Third, it provides new proofs for most activation functions. Fourth, for a given activation and error tolerance, the framework provides precisely the architecture of the corresponding one-hidden neural network with a predetermined number of neurons and the values of weights/biases. Fifth, the framework allows us to abstractly present the first universal approximation with a favorable non-asymptotic rate. Sixth, our framework also provides insights into the developments, and hence providing constructive derivations, of some of the existing approaches.