How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks

How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
复制标题

DOI:
--
复制
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Keyulu Xu;Jingling Li;Mozhi Zhang;S. Du;K. Kawarabayashi;S. Jegelka
Keyulu Xu;Jingling Li;Mozhi Zhang;S. Du;K. Kawarabayashi;S. Jegelka
中科院分区:
其他
文献类型:
--
作者:
Keyulu Xu;Jingling Li;Mozhi Zhang;S. Du;K. Kawarabayashi;S. Jegelka

文献摘要

被引文献

相似文献

我们研究如何通过梯度下降来推断的神经网络,即他们在支持训练分布的支持之外学到的知识。先前的作品报告了用神经网络推断的混合经验结果:而多层感知器(MLP)在某些简单任务中不能很好地推断出具有MLP模块的结构化网络,但在更复杂的任务中显示出一些成功。为了理论解释,我们确定了MLP和GNN良好推断的条件。首先,我们量化了从原始方向迅速收敛到线性函数的观察结果,这意味着Relu MLP不会推断大多数非线性函数。但是,当训练分布足够“多样化”时,他们可以证明他们可以学习线性目标功能。其次,与分析GNN的成功和局限性有关,这些结果提出了一个假设,我们提供了理论和经验证据:GNN在外推算法任务中的成功,将算法任务推出到新数据(例如,较大的图形或边缘)涉及编码任务任务的关系 - 体系结构或功能中的特异性非线性。
We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution. Previous works report mixed empirical results when extrapolating with neural networks: while multilayer perceptrons (MLPs) do not extrapolate well in certain simple tasks, Graph Neural Network (GNN), a structured network with MLP modules, has shown some success in more complex tasks. Working towards a theoretical explanation, we identify conditions under which MLPs and GNNs extrapolate well. First, we quantify the observation that ReLU MLPs quickly converge to linear functions along any direction from the origin, which implies that ReLU MLPs do not extrapolate most non-linear functions. But, they can provably learn a linear target function when the training distribution is sufficiently "diverse". Second, in connection to analyzing successes and limitations of GNNs, these results suggest a hypothesis for which we provide theoretical and empirical evidence: the success of GNNs in extrapolating algorithmic tasks to new data (e.g., larger graphs or edge weights) relies on encoding task-specific non-linearities in the architecture or features.