Bayesian Contextual Bandits for Hyper Parameter Optimization

Bayesian Contextual Bandits for Hyper Parameter Optimization
复制标题

用于超参数优化的贝叶斯上下文老虎机

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
3.9
通讯作者:
Yong Yu
Yong Yu
中科院分区:
计算机科学3区
文献类型:
--
作者:
Guoxin Sui;Yong Yu

文献摘要

被引文献

相似文献

超参数优化(HPO)是现代机器学习系统中的关键步骤。贝叶斯优化(BO)在HPO中表现出了很大的希望,其中参数评估是通过黑盒优化过程进行的。然而,BO的主要缺点在于昂贵的计算成本,这限制了其在大型深度模型中的应用。因此,通过利用迭代训练过程提供的部分信息来减少所需的训练时期的数量是有强烈动机的。最近的进展使用概率模型,显式地推断学习曲线,以提前停止性能较差的配置的训练。然而,这些方法对学习曲线的分布施加了强的先验假设,并且涉及更大的计算复杂度或严重依赖于预定义的规则,这对于不同的情况是不通用的。为了解决在无限参数搜索空间和时间范围内分配训练资源的挑战,我们研究了贝叶斯上下文强盗设置下的HPO问题,并推导出几种信息有效和可扩展的全动态策略。大量的实验表明,该方法可以显着加快HPO过程。
Hyper parameter optimization (HPO) is a crucial step in modern machine learning systems. Bayesian optimization (BO) has shown great promise in HPO, where the parameter evaluation is conducted through a black-box optimization procedure. However, the main drawback of BO lies in the expensive computation cost, which limits its application in large deep models. Hence there is strong motivation to reduce the number of epochs of training required by leveraging the partial information provided by iterative training procedures. Recent advancements use probabilistic models that extrapolate learning curves explicitly to early-stop the training of poor-performing configurations. However, these approaches impose a strong prior assumption on the distribution of learning curves and involve much larger computational complexity or rely heavily on predefined rules, which is not general for different cases. To tackle the challenge of training resource allocation in infinite parameter search space and in time horizon, we study HPO problem in Bayesian contextual bandits setting and derive several fully-dynamic strategies that are information-efficient and scalable. Extensive experiments demonstrate that the proposed method can significantly speed up the HPO process.