Benefits of Depth in Neural Networks

Benefits of Depth in Neural Networks
复制标题

DOI:
--
复制
发表时间:
2016-02
期刊:
--
影响因子:
--
通讯作者:
Matus Telgarsky
Matus Telgarsky
中科院分区:
其他
文献类型:
--
作者:
Matus Telgarsky

文献摘要

被引文献

相似文献

对于任何正整数$k$,存在具有$\Theta(k^3)$层、每层$\Theta(1)$节点和$\Theta(1)$不同参数的神经网络,这些神经网络不能被具有$\mathcal{O}(k)$层的网络近似,除非它们是指数大的--它们必须具有$\Omega(2^k)$节点。在这里,针对称为“半代数门”的一类节点证明了这一结果,其中包括ReLU,最大值,指示符和分段多项式函数的常见选择,因此不仅建立了具有ReLU门的标准网络的深度优势,而且建立了具有ReLU和最大化门的卷积网络,和积网络和提升决策树。(在最后一种情况下,使用更强的分离:需要$\Omega(2^{k^3})$个树节点)。
For any positive integer $k$, there exist neural networks with $\Theta(k^3)$ layers, $\Theta(1)$ nodes per layer, and $\Theta(1)$ distinct parameters which can not be approximated by networks with $\mathcal{O}(k)$ layers unless they are exponentially large --- they must possess $\Omega(2^k)$ nodes. This result is proved here for a class of nodes termed "semi-algebraic gates" which includes the common choices of ReLU, maximum, indicator, and piecewise polynomial functions, therefore establishing benefits of depth against not just standard networks with ReLU gates, but also convolutional networks with ReLU and maximization gates, sum-product networks, and boosted decision trees (in this last case with a stronger separation: $\Omega(2^{k^3})$ total tree nodes are required).