Learning hyperparameters for neural network models using Hamiltonian dynamics

Learning hyperparameters for neural network models using Hamiltonian dynamics
复制标题

使用哈密顿动力学学习神经网络模型的超参数

DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Kiam Choo
Kiam Choo
中科院分区:
--
文献类型:
--
作者:
Kiam Choo

文献摘要

被引文献

相似文献

使用汉密尔顿动力学的神经网络模型学习超参数kiam choo科学科学系多伦多计算机科学系2000年,我们考虑使用控制权重组的超级参数的馈电神经网络模型。鉴于一些训练数据,可以通过使用Gibbs采样的混合蒙特卡洛与混合蒙特卡洛进行交替更新权重,并通过混合蒙特卡洛进行交替更新权重,从而获得了重量和超参数的后部分布。但是,对于具有较大隐藏层的网络而言,此方法变得慢。我们通过将超参数纳入混合蒙特卡洛更新中来解决这个问题。然而,大型高参数后部的状态空间区域很大,概率密度较低,而较小的超参数的区域非常小且密度很高。由于混合蒙特卡洛本来可以在此类区域之间固有的移动良好,因此我们对权重进行重新聚集,以使这两个区域更加兼容,只是受到导致无法计算良好步骤的能力的阻碍。我们的e Orts没有任何改进的结果,但是我们诊断出其原因,并提出了未来的研究方向。 ii奉献精神,我将这一论文奉献给了我的家人,这些年来,他们已经接受了我的流浪。我特别将其奉献给我的母亲,母亲的力量和我会继续下去。致谢我感谢Radford Neal教授对本文的宝贵指导。我还要感谢Faisal Qureshi在本文期间的有益建议和友谊。还要感谢Arti Cial Intelligence实验室的其他居民,他们以某种方式帮助我完成了这一论文。 iii
Learning Hyperparameters for Neural Network Models Using Hamiltonian Dynamics Kiam Choo Master of Science Graduate Department of Computer Science University of Toronto 2000 We consider a feedforward neural network model with hyperparameters controlling groups of weights. Given some training data, the posterior distribution of the weights and the hyperparameters can be obtained by alternately updating the weights with hybrid Monte Carlo and sampling from the hyperparameters using Gibbs sampling. However, this method becomes slow for networks with large hidden layers. We address this problem by incorporating the hyperparameters into the hybrid Monte Carlo update. However, the region of state space under the posterior with large hyperparameters is huge and has low probability density, while the region with small hyperparameters is very small and very high density. As hybrid Monte Carlo inherently does not move well between such regions, we reparameterize the weights to make the two regions more compatible, only to be hampered by the resulting inability to compute good stepsizes. No de nite improvement results from our e orts, but we diagnose the reasons for that, and suggest future directions of research. ii Dedication I dedicate this thesis to my family, who have accepted my wanderings over the years. I especially dedicate this to my mother, whose strength and will I carry on. Acknowledgements I thank Prof. Radford Neal for his invaluable guidance on this thesis. I also thank Faisal Qureshi for his helpful suggestions and friendship during this thesis. Thanks also to the other inhabitants of the Arti cial Intelligence Laboratory who have helped me in some way to complete this thesis. iii