Fast Bayesian hyperparameter optimization on large datasets

Fast Bayesian hyperparameter optimization on large datasets
复制标题

DOI:
10.1214/17-ejs1335si
复制
发表时间:
2017-01-01
影响因子:
1.1
通讯作者:
Hutter, Frank
Hutter, Frank
中科院分区:
数学3区
文献类型:
--
作者:
Klein, Aaron;Falkner, Stefan;Hutter, Frank

文献摘要

被引文献

相似文献

贝叶斯优化已成为机器学习算法超参数优化的成功工具,如支持向量机或深度神经网络。尽管取得了成功,但对于大型数据集,培训和验证单个配置通常需要数小时、数天甚至数周的时间,这限制了可实现的性能。为了加速超参数优化,我们提出了一个作为训练集大小函数的验证误差的生成性模型,该模型在优化过程中学习,并允许通过外推到整个数据集来探索较小子集上的初始配置。我们构造了一个贝叶斯优化过程,称为Fbolas,它将损失和训练时间建模为数据集大小的函数,并自动权衡关于全局最优的高信息增益和计算成本。优化支持向量机和深度神经网络的实验表明,Fbolas找到高质量解的速度通常是其他最先进的贝叶斯优化方法或最近提出的强盗策略Hyperband的10到100倍。
Bayesian optimization has become a successful tool for optimizing the hyperparameters of machine learning algorithms, such as support vector machines or deep neural networks. Despite its success, for large datasets, training and validating a single configuration often takes hours, days, or even weeks, which limits the achievable performance. To accelerate hyperparameter optimization, we propose a generative model for the validation error as a function of training set size, which is learned during the optimization process and allows exploration of preliminary configurations on small subsets, by extrapolating to the full dataset. We construct a Bayesian optimization procedure, dubbed Fabolas, which models loss and training time as a function of dataset size and automatically trades off high information gain about the global optimum against computational cost. Experiments optimizing support vector machines and deep neural networks show that Fabolas often finds high-quality solutions 10 to 100 times faster than other state-of-the-art Bayesian optimization methods or the recently proposed bandit strategy Hyperband.