Bayesian estimation and model averaging of convolutional neural networks by hypernetwork

Bayesian estimation and model averaging of convolutional neural networks by hypernetwork
复制标题

DOI:
10.1587/nolta.10.45
复制
发表时间:
2019
期刊:
Nonlinear Theory and Its Applications, IEICE
影响因子:
--
通讯作者:
K. Ukai;Takashi Matsubara;K. Uehara
K. Ukai;Takashi Matsubara;K. Uehara
中科院分区:
其他
文献类型:
--
作者:
K. Ukai;Takashi Matsubara;K. Uehara

文献摘要

相似文献

神经网络具有学习复杂表征的强大能力,并在各种任务中取得了显著成果。然而,由于训练样本数量有限,它们容易出现过拟合,因此对神经网络的学习过程进行正则化至关重要。在本文中,我们提出一种正则化方法,该方法使用超网络将大型卷积神经网络的参数估计为概率分布,超网络可生成另一个网络的参数。此外,我们进行模型平均以提高网络性能。然后,我们将所提出的方法应用于诸如宽残差网络之类的大型模型。实验结果表明,我们的方法及其模型平均性能优于常用的带L2正则化的最大后验估计。
: Neural networks have a rich ability to learn complex representations and have achieved remarkable results in various tasks. However, they are prone to overfitting owing to the limited number of training samples and regularizing the learning process of neural networks is essential. In this paper, we propose a regularization method that estimates the parameters of a large convolutional neural network as probabilistic distributions using a hypernetwork, which generates the parameters of another network. Additionally, we perform model averaging to improve the network performance. Then, we apply the proposed method to a large model such as wide residual networks. The experimental results demonstrate that our method and its model averaging outperform the commonly used maximum a posteriori estimation with L2 regularization.