Automated Machine Learning on Big Data using Stochastic Algorithm Tuning

Automated Machine Learning on Big Data using Stochastic Algorithm Tuning
复制标题

使用随机算法调整的大数据自动机器学习

DOI:
--
复制
发表时间:
2014
期刊:
影响因子:
--
通讯作者:
S. Roberts
S. Roberts
中科院分区:
--
文献类型:
--
作者:
T. Nickson;Michael A. Osborne;S. Reece;S. Roberts

文献摘要

被引文献

相似文献

我们介绍了一种针对大数据任务自动化机器学习(ML)的方法,通过对ML算法参数和超参数进行可扩展的随机贝叶斯优化。通常情况下,ML算法参数的关键调整依赖于专家的领域专业知识,沿着费力的手工调整,蛮力搜索或冗长的采样运行。在这种背景下,贝叶斯优化在自动化参数调整中的应用越来越多,使得ML算法甚至对非专家也可以访问。然而,贝叶斯优化的现有技术无法扩展到将现实模型拟合到复杂的大数据所需的大量算法性能评估。我们在这里描述了一个随机的,稀疏的,贝叶斯优化策略来解决这个问题,使用成千上万的数据子集上的算法性能的噪声评估,以有效地训练大数据的算法。我们提供了一个全面的基准可能的稀疏化策略贝叶斯优化,得出的结论是Nystrom近似提供了最好的缩放和性能的真实的任务。我们提出的算法在调整真实的大数据上的高斯过程时间序列预测任务的参数方面比现有技术有了很大的改进。
We introduce a means of automating machine learning (ML) for big data tasks, by performing scalable stochastic Bayesian optimisation of ML algorithm parameters and hyper-parameters. More often than not, the critical tuning of ML algorithm parameters has relied on domain expertise from experts, along with laborious hand-tuning, brute search or lengthy sampling runs. Against this background, Bayesian optimisation is finding increasing use in automating parameter tuning, making ML algorithms accessible even to non-experts. However, the state of the art in Bayesian optimisation is incapable of scaling to the large number of evaluations of algorithm performance required to fit realistic models to complex, big data. We here describe a stochastic, sparse, Bayesian optimisation strategy to solve this problem, using many thousands of noisy evaluations of algorithm performance on subsets of data in order to effectively train algorithms for big data. We provide a comprehensive benchmarking of possible sparsification strategies for Bayesian optimisation, concluding that a Nystrom approximation offers the best scaling and performance for real tasks. Our proposed algorithm demonstrates substantial improvement over the state of the art in tuning the parameters of a Gaussian Process time series prediction task on real, big data.