Deep Network Approximation: Achieving Arbitrary Accuracy with Fixed Number of Neurons

Deep Network Approximation: Achieving Arbitrary Accuracy with Fixed Number of Neurons
复制标题

DOI:
--
复制
发表时间:
2021-07
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang
Zuowei Shen;Haizhao Yang;Shijun Zhang
中科院分区:
其他
文献类型:
--
作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang

文献摘要

相似文献

本文开发了简单的前馈神经网络,实现了所有连续函数与固定有限数量的神经元的通用逼近属性。这些神经网络很简单,因为它们是用一个简单的、可计算的、连续的激活函数$\sigma$利用三角波函数和softsign函数设计的。我们首先证明了宽度为36 d(2d+1)$,深度为11 $的$\sigma$激活网络可以在任意小的误差内逼近d$维超立方体上的任意连续函数.因此,对于监督学习及其相关的回归问题,由这些网络生成的大小不小于$36 d(2d+1)\x 11$的假设空间在连续函数空间$C([a,B]^d)$中是稠密的,因此在勒贝格空间$L^p([a,B]^d)$中是稠密的,其中$p\in [1,\infty)$。此外,我们证明了当存在两两不相交的有界闭子集$\mathbb{R}^d$使得同类样本位于同一子集中时,由图像和信号分类产生的分类函数在由$\sigma$激活的网络生成的假设空间中,网络的宽度为$36d(2d+1)$,深度为$12$.最后,我们使用数值实验来证明,用我们的激活函数替换整流线性单元(ReLU)激活函数会改善实验结果。
This paper develops simple feed-forward neural networks that achieve the universal approximation property for all continuous functions with a fixed finite number of neurons. These neural networks are simple because they are designed with a simple, computable, and continuous activation function $\sigma$ leveraging a triangular-wave function and the softsign function. We first prove that $\sigma$-activated networks with width $36d(2d+1)$ and depth $11$ can approximate any continuous function on a $d$-dimensional hypercube within an arbitrarily small error. Hence, for supervised learning and its related regression problems, the hypothesis space generated by these networks with a size not smaller than $36d(2d+1)\times 11$ is dense in the continuous function space $C([a,b]^d)$ and therefore dense in the Lebesgue spaces $L^p([a,b]^d)$ for $p\in [1,\infty)$. Furthermore, we show that classification functions arising from image and signal classification are in the hypothesis space generated by $\sigma$-activated networks with width $36d(2d+1)$ and depth $12$ when there exist pairwise disjoint bounded closed subsets of $\mathbb{R}^d$ such that the samples of the same class are located in the same subset. Finally, we use numerical experimentation to show that replacing the rectified linear unit (ReLU) activation function by ours would improve the experiment results.