Learning hyperparameters for neural network models using Hamiltonian dynamics
Learning hyperparameters for neural network models using Hamiltonian dynamics
复制标题
使用哈密顿动力学学习神经网络模型的超参数
DOI:
--
复制
发表时间:
2000
期刊:
影响因子:
--
通讯作者:
Kiam Choo
中科院分区:
文献类型:
--
作者:
Kiam Choo
Learning Hyperparameters for Neural Network Models Using Hamiltonian Dynamics Kiam Choo Master of Science Graduate Department of Computer Science University of Toronto 2000 We consider a feedforward neural network model with hyperparameters controlling groups of weights. Given some training data, the posterior distribution of the weights and the hyperparameters can be obtained by alternately updating the weights with hybrid Monte Carlo and sampling from the hyperparameters using Gibbs sampling. However, this method becomes slow for networks with large hidden layers. We address this problem by incorporating the hyperparameters into the hybrid Monte Carlo update. However, the region of state space under the posterior with large hyperparameters is huge and has low probability density, while the region with small hyperparameters is very small and very high density. As hybrid Monte Carlo inherently does not move well between such regions, we reparameterize the weights to make the two regions more compatible, only to be hampered by the resulting inability to compute good stepsizes. No de nite improvement results from our e orts, but we diagnose the reasons for that, and suggest future directions of research. ii Dedication I dedicate this thesis to my family, who have accepted my wanderings over the years. I especially dedicate this to my mother, whose strength and will I carry on. Acknowledgements I thank Prof. Radford Neal for his invaluable guidance on this thesis. I also thank Faisal Qureshi for his helpful suggestions and friendship during this thesis. Thanks also to the other inhabitants of the Arti cial Intelligence Laboratory who have helped me in some way to complete this thesis. iii