Deep Neural Networks Learn Non-Smooth Functions Effectively

Deep Neural Networks Learn Non-Smooth Functions Effectively
复制标题

DOI:
--
复制
发表时间:
2018-02
期刊:
--
影响因子:
--
通讯作者:
M. Imaizumi;K. Fukumizu
M. Imaizumi;K. Fukumizu
中科院分区:
其他
文献类型:
--
作者:
M. Imaizumi;K. Fukumizu

文献摘要

被引文献

相似文献

我们从理论上讨论了为什么深度神经网络(DNN)在某些情况下比其他模型表现得更好,通过研究DNN对非光滑函数的统计特性。虽然DNN在经验上显示出比其他标准方法更高的性能,但理解其机制仍然是一个具有挑战性的问题。从统计理论的一个方面来看,已知许多标准方法获得最佳收敛率,因此很难找到DNN的理论优势。本文填补了这一空白,考虑学习的一类非光滑函数,这是没有涵盖的以前的理论。我们推导了带有ReLU激活的DNN估计量的收敛速度,并表明DNN估计量几乎是估计非光滑函数的最佳估计量,而一些流行的模型没有达到最佳速度。此外,我们的理论结果提供了指导方针,选择适当的层数和边缘的DNN。我们提供了数值实验来支持理论结果。
We theoretically discuss why deep neural networks (DNNs) performs better than other models in some cases by investigating statistical properties of DNNs for non-smooth functions. While DNNs have empirically shown higher performance than other standard methods, understanding its mechanism is still a challenging problem. From an aspect of the statistical theory, it is known many standard methods attain optimal convergence rates, and thus it has been difficult to find theoretical advantages of DNNs. This paper fills this gap by considering learning of a certain class of non-smooth functions, which was not covered by the previous theory. We derive convergence rates of estimators by DNNs with a ReLU activation, and show that the estimators by DNNs are almost optimal to estimate the non-smooth functions, while some of the popular models do not attain the optimal rate. In addition, our theoretical result provides guidelines for selecting an appropriate number of layers and edges of DNNs. We provide numerical experiments to support the theoretical results.