Optimal Approximation Rate of ReLU Networks in terms of Width and Depth
Optimal Approximation Rate of ReLU Networks in terms of Width and Depth
复制标题
DOI:
10.1016/j.matpur.2021.07.009
复制
发表时间:
2021-02
期刊:
影响因子:
--
通讯作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang
中科院分区:
文献类型:
--
作者:
Zuowei Shen;Haizhao Yang;Shijun Zhang
This paper concentrates on the approximation power of deep feed-forward neural networks in terms of width and depth. It is proved by construction that ReLU networks with width O (max{d⌊ N 1/d⌋, N+ 2}) and depth O (L) can approximate a Hölder continuous function on [0, 1] d with an approximation rate O (λ d (N 2 L 2 ln N)− α/d), where α∈(0, 1] and λ> 0 are Hölder order and constant, respectively. Such a rate is optimal up to a constant in terms of width and depth separately, while existing results are only nearly optimal without the logarithmic factor in the approximation rate. More generally, for an arbitrary continuous function f on [0, 1] d, the approximation rate becomes O (d ω f ((N 2 L 2 ln N)− 1/d)), where ω f (⋅) is the modulus of continuity. We also extend our analysis to any continuous function f on a bounded set. Particularly, if ReLU networks with depth 31 and width O (N) are used to approximate one-dimensional Lipschitz continuous functions on [0, 1] with a Lipschitz constant λ> 0, the approximation rate in terms of the total number of parameters, W= O (N 2), becomes O (λ W ln W), which has not been discovered in the literature for fixed-depth ReLU networks.