Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis

Deep Architecture Connectivity Matters for Its Convergence: A Fine-Grained Analysis
复制标题

DOI:
10.48550/arxiv.2205.05662
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Wuyang Chen;Wei Huang;Xinyu Gong;B. Hanin;Zhangyang Wang
Wuyang Chen;Wei Huang;Xinyu Gong;B. Hanin;Zhangyang Wang
中科院分区:
其他
文献类型:
--
作者:
Wuyang Chen;Wei Huang;Xinyu Gong;B. Hanin;Zhangyang Wang

文献摘要

相似文献

由人类或自动机器学习(AutoML)算法设计的先进深度神经网络(DNN)正变得日益复杂。不同的操作通过复杂的连接模式相连,例如各种类型的跳跃连接。这些拓扑结构在经验上是有效的,并且据观察通常能使损失曲面更平滑并促进梯度流动。然而,要从原理上理解它们对DNN容量或可训练性的影响,以及理解为什么一种特定的连接模式在某个方面优于另一种模式,仍然很困难。在这项工作中,我们从理论上精细地描述了连接模式在梯度下降训练下对DNN收敛性的影响。通过分析一个宽网络的神经网络高斯过程(NNGP),我们能够描绘出NNGP核的频谱如何通过一种特定的连接模式传播,以及这如何影响收敛速度的界限。作为我们研究结果的一个实际应用,我们表明通过对“没有前景”的连接模式进行简单筛选,我们可以减少要评估的模型数量,并在没有任何额外开销的情况下显著加速大规模神经架构搜索。代码可在以下网址获取:https://github.com/VITA - Group/architecture_convergence
Advanced deep neural networks (DNNs), designed by either human or AutoML algorithms, are growing increasingly complex. Diverse operations are connected by complicated connectivity patterns, e.g., various types of skip connections. Those topological compositions are empirically effective and observed to smooth the loss landscape and facilitate the gradient flow in general. However, it remains elusive to derive any principled understanding of their effects on the DNN capacity or trainability, and to understand why or in which aspect one specific connectivity pattern is better than another. In this work, we theoretically characterize the impact of connectivity patterns on the convergence of DNNs under gradient descent training in fine granularity. By analyzing a wide network's Neural Network Gaussian Process (NNGP), we are able to depict how the spectrum of an NNGP kernel propagates through a particular connectivity pattern, and how that affects the bound of convergence rates. As one practical implication of our results, we show that by a simple filtration on"unpromising"connectivity patterns, we can trim down the number of models to evaluate, and significantly accelerate the large-scale neural architecture search without any overhead. Code is available at: https://github.com/VITA-Group/architecture_convergence.