Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes

Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
影响因子:
3.9
通讯作者:
Roman Novak;Lechao Xiao;Yasaman Bahri;Jaehoon Lee;Greg Yang;Jiri Hron;Daniel A. Abolafia;Jeffrey Pennin
Roman Novak;Lechao Xiao;Yasaman Bahri;Jaehoon Lee;Greg Yang;Jiri Hron;Daniel A. Abolafia;Jeffrey Pennin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Roman Novak;Lechao Xiao;Yasaman Bahri;Jaehoon Lee;Greg Yang;Jiri Hron;Daniel A. Abolafia;Jeffrey Pennin

文献摘要

被引文献

相似文献

在宽的完全连接的神经网络(FCN)和高斯过程(GPS)之间存在先前确定的等效性。例如,这种等效性使得能够由完全贝叶斯,无限宽的训练FCN产生的测试集预测,而无需实例化FCN,而是通过评估相应的GP。在这项工作中,我们得出了具有和没有合并层的多层卷积神经网络(CNN)的类似等效性,并在没有可训练的内核的GPS上实现了CIFAR10的最新情况。我们还引入了一种蒙特卡洛方法,以估计与给定神经网络结构相对应的GP,即使在分析形式的术语太多以至于无法计算可行的情况下。令人惊讶的是,在没有合并层的情况下,与具有重量共享的CNN相对应的GP是相同的。结果,在有限的通道CNN中有益的翻译均值,可以保证接受随机梯度下降(SGD)的训练,在贝叶斯对无限通道限制的贝叶斯待遇中不起作用 - 在两个制度之间不存在的质量差异FCN案件。我们通过实验确认,尽管在某些情况下,经过SGD训练有限的CNN的性能,随着频道计数的增加,相应的GPS的性能接近相应的GPS,但仔细调整SGD训练的CNN可以显着胜过相应的GPS,这表明与SGD训练相比的优势相比,完全贝叶斯参数估计。
There is a previously identified equivalence between wide fully connected neural networks (FCNs) and Gaussian processes (GPs). This equivalence enables, for instance, test set predictions that would have resulted from a fully Bayesian, infinitely wide trained FCN to be computed without ever instantiating the FCN, but by instead evaluating the corresponding GP. In this work, we derive an analogous equivalence for multi-layer convolutional neural networks (CNNs) both with and without pooling layers, and achieve state of the art results on CIFAR10 for GPs without trainable kernels. We also introduce a Monte Carlo method to estimate the GP corresponding to a given neural network architecture, even in cases where the analytic form has too many terms to be computationally feasible. Surprisingly, in the absence of pooling layers, the GPs corresponding to CNNs with and without weight sharing are identical. As a consequence, translation equivariance, beneficial in finite channel CNNs trained with stochastic gradient descent (SGD), is guaranteed to play no role in the Bayesian treatment of the infinite channel limit - a qualitative difference between the two regimes that is not present in the FCN case. We confirm experimentally, that while in some scenarios the performance of SGD-trained finite CNNs approaches that of the corresponding GPs as the channel count increases, with careful tuning SGD-trained CNNs can significantly outperform their corresponding GPs, suggesting advantages from SGD training compared to fully Bayesian parameter estimation.