Effective Minkowski Dimension of Deep Nonparametric Regression: Function Approximation and Statistical Theories

Effective Minkowski Dimension of Deep Nonparametric Regression: Function Approximation and Statistical Theories
复制标题

DOI:
10.48550/arxiv.2306.14859
复制
发表时间:
2023-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Zixuan Zhang;Minshuo Chen;Mengdi Wang;Wenjing Liao;Tuo Zhao
Zixuan Zhang;Minshuo Chen;Mengdi Wang;Wenjing Liao;Tuo Zhao
中科院分区:
其他
文献类型:
--
作者:
Zixuan Zhang;Minshuo Chen;Mengdi Wang;Wenjing Liao;Tuo Zhao

文献摘要

相似文献

现有的深度非参数回归理论表明,当输入数据位于低维流形上时,深度神经网络可以适应内在的数据结构。在真实的世界的应用中,这样的数据恰好位于低维流形上的假设是严格的。本文引入了一个宽松的假设,即输入数据集中在$\mathcal{S}$表示的$\mathbb{R}^d$的一个子集周围,$\mathcal{S}$的内禀维数可以用一种新的复杂度表示法--有效Minkowski维数来刻画。我们证明了深度非参数回归的样本复杂度只依赖于$\mathcal{S}$的有效Minkowski维数,记为$p$。我们进一步说明我们的理论研究结果,考虑非参数回归与各向异性高斯随机设计$N(0,\Sigma)$,其中$\Sigma$是满秩。当$\Sigma$的特征值有指数或多项式衰减时,这种高斯随机设计的有效Minkowski维数分别为$p=\mathcal{O}(\sqrt{\log n})$或$p=\mathcal{O}(n^\gamma)$,其中$n$是样本大小,$\gamma\in(0,1)$是取决于多项式衰减率的小常数。我们的理论表明,当流形假设不成立时,深度神经网络仍然可以适应数据的有效Minkowski维数,并在中等样本量下规避环境维数的灾难。
Existing theories on deep nonparametric regression have shown that when the input data lie on a low-dimensional manifold, deep neural networks can adapt to the intrinsic data structures. In real world applications, such an assumption of data lying exactly on a low dimensional manifold is stringent. This paper introduces a relaxed assumption that the input data are concentrated around a subset of $\mathbb{R}^d$ denoted by $\mathcal{S}$, and the intrinsic dimension of $\mathcal{S}$ can be characterized by a new complexity notation -- effective Minkowski dimension. We prove that, the sample complexity of deep nonparametric regression only depends on the effective Minkowski dimension of $\mathcal{S}$ denoted by $p$. We further illustrate our theoretical findings by considering nonparametric regression with an anisotropic Gaussian random design $N(0,\Sigma)$, where $\Sigma$ is full rank. When the eigenvalues of $\Sigma$ have an exponential or polynomial decay, the effective Minkowski dimension of such an Gaussian random design is $p=\mathcal{O}(\sqrt{\log n})$ or $p=\mathcal{O}(n^\gamma)$, respectively, where $n$ is the sample size and $\gamma\in(0,1)$ is a small constant depending on the polynomial decay rate. Our theory shows that, when the manifold assumption does not hold, deep neural networks can still adapt to the effective Minkowski dimension of the data, and circumvent the curse of the ambient dimensionality for moderate sample sizes.