On the universality of deep learning

On the universality of deep learning
复制标题

DOI:
--
复制
发表时间:
2020
期刊:
ArXiv
影响因子:
--
通讯作者:
E. Abbe;Colin Sandon
E. Abbe;Colin Sandon
中科院分区:
其他
文献类型:
--
作者:
E. Abbe;Colin Sandon

文献摘要

被引文献

相似文献

本文表明,深度学习,即,用SGD训练的神经网络,可以在多时间内学习任何可以通过某种算法在多时间内学习的函数类,包括奇偶校验。这个普遍的结果被进一步证明是稳健的,即,它在梯度上可能存在多噪声的情况下成立,这给出了深度学习和统计查询算法之间的分离,因为后者由于像奇偶校验这样的情况而不是通用的。这也表明,基于SGD的深度学习不会受到Minsky-Papert '69讨论的感知器的限制。本文进一步补充了这一结果与下降算法的泛化误差的下限,这意味着特别是强大的普适性崩溃,如果梯度是平均足够大的批次的样本在全GD,而不是更少的样本在SGD。
This paper shows that deep learning, i.e., neural networks trained by SGD, can learn in polytime any function class that can be learned in polytime by some algorithmm, including parities. This universal result is further shown to be robust, i.e., it holds under possibly poly-noise on the gradients, which gives a separation between deep learning and statistical query algorithms, as the latter are not comparably universal due to cases like parities. This also shows that SGD-based deep learning does not suffer from the limitations of the perceptron discussed by Minsky-Papert ’69. The paper further complement this result with a lower-bound on the generalization error of descent algorithms, which implies in particular that the robust universality breaks down if the gradients are averaged over large enough batches of samples as in full-GD, rather than fewer samples as in SGD.