Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth

Deep Network with Approximation Error Being Reciprocal of Width to Power of Square Root of Depth
复制标题

DOI:
10.1162/neco_a_01364
复制
发表时间:
2020-06
期刊:
影响因子:
2.9
通讯作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang
Zuowei Shen;Haizhao Yang;Shijun Zhang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang

文献摘要

被引文献

相似文献

引入了具有超逼近能力的新网络。该网络是在每个神经元中使用 Floor (⌊x⌋) 或 ReLU (max{0,x}) 激活函数构建的;因此,我们将这种网络称为 Floor-ReLU 网络。对于任何超参数 NεN+ 和 LεN+,我们表明宽度 max{d,5N+13} 和深度 64dL+3 的 Floor-ReLU 网络可以均匀地近似 [0,1]d 上的 Hölder 函数 f,近似误差为 3λdα/2N-αL,其中 αε(0,1] 和 λ 分别是 Hölder 阶数和常数。更一般地,对于任意连续函数 f [0,1]d 具有连续性模 ωf(·),相长近似率为 ωf(dN-L)+2ωf(d)N-L 因此,当 ωf(r) 随着 r→0 的变化适中时(例如,对于 Hölder 连续函数,ωf(r)≲rα),这种新型网络克服了近似幂的维数诅咒,因为在我们的近似率本质上是 d 乘以 N 和 L 的函数,与连续模内的 d 无关。
A new network with super-approximation power is introduced. This network is built with Floor (⌊x⌋) or ReLU (max{0,x}) activation function in each neuron; hence, we call such networks Floor-ReLU networks. For any hyperparameters N∈N+ and L∈N+, we show that Floor-ReLU networks with width max{d,5N+13} and depth 64dL+3 can uniformly approximate a Hölder function f on [0,1]d with an approximation error 3λdα/2N-αL, where α∈(0,1] and λ are the Hölder order and constant, respectively. More generally for an arbitrary continuous function f on [0,1]d with a modulus of continuity ωf(·), the constructive approximation rate is ωf(dN-L)+2ωf(d)N-L. As a consequence, this new class of networks overcomes the curse of dimensionality in approximation power when the variation of ωf(r) as r→0 is moderate (e.g., ωf(r)≲rα for Hölder continuous functions), since the major term to be considered in our approximation rate is essentially d times a function of N and L independent of d within the modulus of continuity.