Gradient conjugate priors and multi-layer neural networks

Gradient conjugate priors and multi-layer neural networks
复制标题

DOI:
10.1016/j.artint.2019.103184
复制
发表时间:
2020-01-01
影响因子:
14.4
通讯作者:
Stuke, Hannes
Stuke, Hannes
中科院分区:
计算机科学2区
文献类型:
--
作者:
Gurevich, Pavel;Stuke, Hannes

文献摘要

被引文献

相似文献

本文涉及通过人工神经网络学习观测数据的概率分布。我们建议适用于神经网络的所谓梯度共轭先验(GCP)更新,它是共轭先验的经典贝叶斯更新的修改。我们在梯度共轭先验更新和预测分布的对数似然最大化之间建立了联系。与贝叶斯神经网络不同,我们使用神经网络的确定性权重,而是假设地面真值分布是具有未知均值和方差的正态分布,并通过神经网络学习这些未知均值和方差的先验(正态伽玛分布)的参数。参数的更新是使用梯度完成的,该梯度在每一步都指向最小化与先验分布和后验分布(均为正态伽玛)的 Kullback-Leibler 散度。我们获得先验参数的相应动力系统并分析其性质。特别是,我们研究了所有先验参数的限制行为,并展示了它与经典完整贝叶斯更新的情况有何不同。结果在合成数据集和现实世界数据集上得到验证。 (C) 2019 Elsevier B.V. 保留所有权利。
The paper deals with learning probability distributions of observed data by artificial neural networks. We suggest a so-called gradient conjugate prior (GCP) update appropriate for neural networks, which is a modification of the classical Bayesian update for conjugate priors. We establish a connection between the gradient conjugate prior update and the maximization of the log-likelihood of the predictive distribution. Unlike for the Bayesian neural networks, we use deterministic weights of neural networks, but rather assume that the ground truth distribution is normal with unknown mean and variance and learn by the neural networks the parameters of a prior (normal-gamma distribution) for these unknown mean and variance. The update of the parameters is done, using the gradient that, at each step, directs towards minimizing the Kullback-Leibler divergence from the prior to the posterior distribution (both being normal-gamma). We obtain a corresponding dynamical system for the prior's parameters and analyze its properties. In particular, we study the limiting behavior of all the prior's parameters and show how it differs from the case of the classical full Bayesian update. The results are validated on synthetic and real world data sets. (C) 2019 Elsevier B.V. All rights reserved.