Sherpa: Hyperparameter Optimization for Machine Learning Models

Sherpa: Hyperparameter Optimization for Machine Learning Models
复制标题

Sherpa:机器学习模型的超参数优化

DOI:
--
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
P. Baldi
P. Baldi
中科院分区:
--
文献类型:
--
作者:
L. Hertel;Julian Collado;Peter Sadowski;P. Baldi

文献摘要

参考文献

被引文献

相似文献

Sherpa是一个免费的开源超参数优化库,用于机器学习模型。它是为计算量大的迭代函数求值问题而设计的,如深度神经网络的超参数整定。有了Sherpa,科学家可以使用各种功能强大且可互换的算法快速优化超参数。此外,该框架使实现自定义算法变得容易。Sherpa既可以在单台机器上运行,也可以通过网格调度器在集群上运行,只需极少的配置。最后,交互式仪表板使用户能够在模型训练时查看模型的进度,取消试验,并探索哪种超参数组合效果最好。Sherpa通过自动化模型调优的繁琐方面并为开发自动化超参数调优策略提供可扩展框架,从而增强了机器学习研究人员的能力。其源代码和文档分别可从https://github.com/LarsHH/sherpa和https://parameter-sherpa.readthedocs.io/获得。可以在https://youtu.be/L95sasMLgP4上找到一个演示。机器学习模型的超参数优化算法已经在诸如Spearmint[15]、HyperOpt[2]、Auto-Weka 2.0[9]和谷歌Vizier[5]等软件包中实现。Spearmint是一个基于贝叶斯优化的Python库,使用高斯过程。使用标记语言YAML指定超参数探索值,并通过SGE和MongoDB在网格上运行。总的来说,它结合了贝叶斯优化和分布式训练的能力。HyperOpt是一个使用MongoDB实现并行计算的超参数优化框架。用户手动启动从HyperOpt实例接收任务的worker。它提供了基于Parzen估计器树的随机搜索和贝叶斯优化的使用。Auto-WEKA 2.0在WEKA机器学习框架内实现了SMAC[6]算法的自动模型选择和超参数优化。它提供了一个图形用户界面,并支持在一台机器上并行运行。它旨在为新手用户提供可访问性,并专门针对选择模型的问题。Auto-WEKA与Auto-Sklearn[4]和Auto-Net[11]相关,它们专门专注于调整Scikit-Learn模型和完全连接的第32届神经信息处理系统会议(NIPS 2018),加拿大montracimal。表1:与现有库的比较:Spearmint Auto-WEKA HyperOpt谷歌Vizier Sherpa Early stop No No No Yes Dashboard/GUI Yes Yes No Yes分布式Yes No Yes Yes开源Yes Yes No Yes #算法2 1 2 3 5神经网络分别在Lasagne中。Auto-WEKA、Auto-Sklearn和Auto-Net侧重于端到端的自动化方法。这对新手来说很容易,但将用户限制在各自的机器学习库和它实现的模型中。相比之下,我们的工作旨在为用户在库、模型和超参数优化算法选择上提供更大的灵活性。谷歌Vizier是谷歌为其云机器学习平台提供的一项服务。它结合了贝叶斯优化的最新创新,如迁移学习,并通过仪表板提供可视化。谷歌Vizier为谷歌Cloud用户和谷歌工程师提供了当前超参数优化工具的许多关键功能,但在开源版本中不可用。类似的情况也发生在其他基于云的平台上,比如Microsoft Azure的超参数优化1和Amazon SageMaker的超参数优化2。近年来,机器学习领域经历了巨大的增长。访问开源机器学习库,如Scikit-Learn[14]、Keras[3]、Tensorflow[1]、PyTorch[13]和Caffe[8],允许机器学习研究被社区广泛复制,使从业者可以轻松地将最先进的方法应用于现实世界的问题。机器学习的超参数优化领域最近也出现了许多创新,如Hyperband [10], Population Based Training [7], Neural Architecture Search[17],以及贝叶斯优化的创新,如[16]。虽然其中一些算法的基本实现可能是微不足道的,但以分布式方式评估试验并跟踪结果变得繁琐,这使得用户难以将这些算法应用于实际问题。简而言之,Sherpa旨在策划这些算法的实现,同时提供以分布式方式运行这些算法的基础设施。其目的是使该平台能够从笔记本电脑上的使用扩展到计算网格。
Sherpa is a free open-source hyperparameter optimization library for machine learning models. It is designed for problems with computationally expensive iterative function evaluations, such as the hyperparameter tuning of deep neural networks. With Sherpa, scientists can quickly optimize hyperparameters using a variety of powerful and interchangeable algorithms. Additionally, the framework makes it easy to implement custom algorithms. Sherpa can be run on either a single machine or a cluster via a grid scheduler with minimal configuration. Finally, an interactive dashboard enables users to view the progress of models as they are trained, cancel trials, and explore which hyperparameter combinations are working best. Sherpa empowers machine learning researchers by automating the tedious aspects of model tuning and providing an extensible framework for developing automated hyperparameter-tuning strategies. Its source code and documentation are available at https://github.com/LarsHH/sherpa and https://parameter-sherpa.readthedocs.io/, respectively. A demo can be found at https://youtu.be/L95sasMLgP4. 1 Existing Hyperparameter Optimization Libraries Hyperparameter optimization algorithms for machine learning models have previously been implemented in software packages such as Spearmint [15], HyperOpt [2], Auto-Weka 2.0 [9], and Google Vizier [5] among others. Spearmint is a Python library based on Bayesian optimization using a Gaussian process. Hyperparameter exploration values are specified using the markup language YAML and run on a grid via SGE and MongoDB. Overall, it combines Bayesian optimization with the ability for distributed training. HyperOpt is a hyperparameter optimization framework that uses MongoDB to allow parallel computation. The user manually starts workers which receive tasks from the HyperOpt instance. It offers the use of Random Search and Bayesian optimization based on a Tree of Parzen Estimators. Auto-WEKA 2.0 implements the SMAC [6] algorithm for automatic model selection and hyperparameter optimization within the WEKA machine learning framework. It provides a graphical user interface and supports parallel runs on a single machine. It is meant to be accessible for novice users and specifically targets the problem of choosing a model. Auto-WEKA is related to Auto-Sklearn [4] and Auto-Net [11] which specifically focus on tuning Scikit-Learn models and fully-connected 32nd Conference on Neural Information Processing Systems (NIPS 2018), Montréal, Canada. Table 1: Comparison to Existing Libraries Spearmint Auto-WEKA HyperOpt Google Vizier Sherpa Early Stopping No No No Yes Yes Dashboard/GUI Yes Yes No Yes Yes Distributed Yes No Yes Yes Yes Open Source Yes Yes Yes No Yes # of Algorithms 2 1 2 3 5 neural networks in Lasagne, respectively. Auto-WEKA, Auto-Sklearn, and Auto-Net focus on an end-to-end automatic approach. This makes it easy for novice users, but restricts the user to the respective machine learning library and the models it implements. In contrast our work aims to give the user more flexibility over library, model and hyper-parameter optimization algorithm selection. Google Vizier is a service provided by Google for its cloud machine learning platform. It incorporates recent innovation in Bayesian optimization such as transfer learning and provides visualizations via a dashboard. Google Vizier provides many key features of a current hyperparameter optimization tool to Google Cloud users and Google engineers, but is not available in an open source version. A similar situation occurs with other cloud based platforms like Microsoft Azure Hyperparameter Tuning 1 and Amazon SageMaker’s Hyperparameter Optimization 2. 2 Need for a new library The field of machine learning has experienced massive growth over recent years. Access to open source machine learning libraries such as Scikit-Learn [14], Keras [3], Tensorflow [1], PyTorch [13], and Caffe [8] allowed research in machine learning to be widely reproduced by the community making it easy for practitioners to apply state of the art methods to real world problems. The field of hyperparameter optimization for machine learning has also seen many innovations recently such as Hyperband [10], Population Based Training [7], Neural Architecture Search [17], and innovation in Bayesian optimization such as [16]. While the basic implementation of some of these algorithms can be trivial, evaluating trials in a distributed fashion and keeping track of results becomes cumbersome which makes it difficult for users to apply these algorithms to real problems. In short, Sherpa aims to curate implementations of these algorithms while providing infrastructure to run these in a distributed way. The aim is for the platform to be scalable from usage on a laptop to a computation grid.
DOI: --
发表时间: 2016-11
期刊: ArXiv
影响因子: --
作者:
Barret Zoph;Quoc V. Le
通讯作者: Barret Zoph;Quoc V. Le