Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality

Adaptivity of deep ReLU network for learning in Besov and mixed smooth Besov spaces: optimal rate and curse of dimensionality
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Taiji Suzuki
Taiji Suzuki
中科院分区:
其他
文献类型:
--
作者:
Taiji Suzuki

文献摘要

被引文献

相似文献

深度学习表明,从视觉识别到自然语言处理的各种任务中的表现都很高,这表明了深度学习的卓越灵活性和适应性。从理论上讲,我们对深度学习进行了新​​的近似和估计误差分析,并在besov空间中的功能及其具有混合光滑度的变体进行了深度学习的新近似和估计误差分析。 BESOV空间是一个相当普遍的功能空间,包括持有人空间和Sobolev空间,尤其可以捕获平滑度的空间不均匀性。通过在BESOV空间中的分析,可以表明深度学习可以达到最小的最佳速率,并胜过任何非自适应(线性)估计量(例如内核脊回归),这表明深度学习对对空间不均匀性具有更高的适应性与其他估计值(例如线性估计值)相比。除此之外,还表明,如果目标函数在混合光滑的BESOV空间中,则深度学习可以避免维度的诅咒。我们还表明,由于其最小值最佳性,收敛速率对维度的依赖性很紧。这些结果支持深度学习的高适应性及其作为特征提取器的优越能力。
Deep learning has shown high performances in various types of tasks from visual recognition to natural language processing, which indicates superior flexibility and adaptivity of deep learning. To understand this phenomenon theoretically, we develop a new approximation and estimation error analysis of deep learning with the ReLU activation for functions in a Besov space and its variant with mixed smoothness. The Besov space is a considerably general function space including the Holder space and Sobolev space, and especially can capture spatial inhomogeneity of smoothness. Through the analysis in the Besov space, it is shown that deep learning can achieve the minimax optimal rate and outperform any non-adaptive (linear) estimator such as kernel ridge regression, which shows that deep learning has higher adaptivity to the spatial inhomogeneity of the target function than other estimators such as linear ones. In addition to this, it is shown that deep learning can avoid the curse of dimensionality if the target function is in a mixed smooth Besov space. We also show that the dependency of the convergence rate on the dimensionality is tight due to its minimax optimality. These results support high adaptivity of deep learning and its superior ability as a feature extractor.