Gaussian Process Behaviour in Wide Deep Neural Networks

Gaussian Process Behaviour in Wide Deep Neural Networks
复制标题

DOI:
--
复制
发表时间:
2018-02
期刊:
ArXiv
影响因子:
--
通讯作者:
A. G. Matthews;Mark Rowland;Jiri Hron;Richard E. Turner;Zoubin Ghahramani
A. G. Matthews;Mark Rowland;Jiri Hron;Richard E. Turner;Zoubin Ghahramani
中科院分区:
其他
文献类型:
--
作者:
A. G. Matthews;Mark Rowland;Jiri Hron;Richard E. Turner;Zoubin Ghahramani

文献摘要

被引文献

相似文献

尽管深层神经网络已经显示出巨大的经验成功,但仍有很多工作要做,以了解其理论特性。在本文中,我们研究了具有以上隐藏层的随机,宽,​​完全连接,前馈网络与具有递归核定义的高斯过程之间的关系。我们表明,在广泛的条件下,随着我们使体系结构越来越宽,隐含的随机函数在分布中收敛到高斯流程,从而将Neal(1996)的现有结果形式化并扩展到了深网。为了从经验上评估收敛速率,我们使用最大的平均差异。然后,我们将有限的贝叶斯深网从文献与高斯流程进行了比较,从关键的预测量角度来看,发现在某些情况下,该协议可能非常接近。我们讨论了高斯过程行为的可取性,并回顾了文献中的非高斯替代模型。
Whilst deep neural networks have shown great empirical success, there is still much work to be done to understand their theoretical properties. In this paper, we study the relationship between random, wide, fully connected, feedforward networks with more than one hidden layer and Gaussian processes with a recursive kernel definition. We show that, under broad conditions, as we make the architecture increasingly wide, the implied random function converges in distribution to a Gaussian process, formalising and extending existing results by Neal (1996) to deep networks. To evaluate convergence rates empirically, we use maximum mean discrepancy. We then compare finite Bayesian deep networks from the literature to Gaussian processes in terms of the key predictive quantities of interest, finding that in some cases the agreement can be very close. We discuss the desirability of Gaussian process behaviour and review non-Gaussian alternative models from the literature.