Exploring Efficient Automated Design Choices for Robust Machine Learning Algorithms
Exploring Efficient Automated Design Choices for Robust Machine Learning Algorithms
批准号:
2748823
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
你熟悉机器学习吗?你有分析和传播各种结果信息的能力吗?您是否愿意协助GCHQ开发和设计新的算法,以实现时间效率和能源效率的解决方案?你是否热衷于开发新颖的方法和使用现代计算架构,使深度学习和高斯过程应用于实际问题变得容易?应用机器学习(ML)目前需要数据科学家做出设计选择。例如,这些选择可能涉及:选择深度神经网络的层数和每层神经元的数量;选择在高斯过程中使用哪个核族。由于机器学习算法通常涉及耗时的训练制度,数据科学家经常发现在(重新)识别候选设计选择和(重新)训练机器学习算法之间进行迭代是很费力的。此外,不同的设计选择可以改变需要考虑的超参数(例如神经元权重或核宽度和交叉协方差项)的数量,也可以改变优化ML算法的超参数的挑战性。由于实践者对这些参数进行敏感性分析的时间有限,设计选择通常是基于估计的性能(作为测试集的平均值计算),如果有的话,对估计中的方差的考虑非常有限。考虑这种差异是很重要的,因为它将决定在操作部署算法时,测试集上的性能准确预测经验性能的可能性有多大。事实上,稳健的性能要求我们不优化超参数(例如,使用随机梯度下降),而是为与数据一致的超参数生成一组样本,然后对这些超参数的采样值进行平均。现有的数值贝叶斯算法可以探索设计选择和与每个设计选择相关的可能的超参数值。这些算法存在成熟的变体,涉及使用马尔可夫链蒙特卡罗(MCMC),可逆跳跃MCMC (RJMCMC)是一种适用于设计选择改变需要考虑的超参数数量的上下文的变体。一般来说,特别是在RJMCMC的情况下,这些成熟的算法是足够缓慢和计算要求,他们被广泛认为是不切实际的,在现实世界的实际使用场景。利物浦大学的最新进展表明,顺序蒙特卡罗(SMC)采样器是数值贝叶斯算法的另一种选择,它有可能提高MCMC算法的时间效率和能源效率。在这种情况下,SMC采样器可以被认为是由一组子算法组成的,这些子算法合作探索设计选择和相关超参数的空间。通过将子算法分布在并行计算资源上,SMC采样器可以提高时间效率。由于子算法只需要避免一次所有失败,因此它们在探索过程中比单个MCMC算法更具冒险精神:这可以带来能源效率的提高。也许令人惊讶的是,SMC采样器自动化设计选择的潜力,同时也探索相关的超参数值,在很大程度上尚未被探索。本博士将研究在这种情况下应用SMC采样器的重大潜力。
英文摘要
Are you familiar with machine learning? Do you have an aptitude for analysing and disseminating information from a variety of outcomes? Would you like to assist GCHQ develop and design new algorithms for both time-efficiency and energy-efficiency solutions? Are you keen on developing novel approaches and using modern computing architectures that make it easy to apply Deep Learning and Gaussian Processes to real problems?Applying Machine Learning (ML) currently requires the data scientist to make design choices. These choices might relate, for example, to: choosing the number of layers and the number of neurons in each layer of a Deep Neural Network; choosing which kernel family to use in a Gaussian Process. Since ML algorithms often involve time-consuming training regimes, data scientists often find it laborious to iterate between (re)-identifying candidate design choices and (re)-training the ML algorithms. Furthermore, different design choices can alter both how many hyper-parameters (e.g. neuron weights or kernel widths and cross-covariance terms) need to considered but also how challenging it is to optimise the hyper-parameters of the ML algorithm. Since practitioners have limited time to perform sensitivity analyses with respect to these parameters, design choice are typically based on estimated performance (calculated as an average over the test set) with very limited, if any, consideration for the variance in this estimate. It is important that this variance is considered since it will determine how likely it is that performance on the test set will accurately predict empirical performance when the algorithm is deployed operationally. Indeed, robust performance requires that we do not optimise the hyper-parameters (e.g. using stochastic gradient descent) but generate a set of samples for the hyper-parameters that are consistent with the data and then average across these sampled values for the hyper-parameters.Numerical Bayesian algorithms exist that can explore the design choices and the possible hyper-parameter values associated with each design choice. Mature variants of these algorithms exist and involve the use of Markov-Chain Monte Carlo (MCMC), with Reversible Jump MCMC (RJMCMC) being a variant applicable in contexts where the design choice alters the number of hyper-parameters that need to be considered. In general, and particularly in the case of RJMCMC, these mature algorithms are sufficiently slow and computationally demanding that they are widely assumed to be impractical for practical use in real-world scenarios.Recent advances at the University of Liverpool have identified that Sequential Monte Carlo (SMC) samplers are an alternative family of numerical Bayesian algorithms that offer the potential to improve on both the time-efficiency and energy-efficiency of MCMC algorithms. In this context, SMC samplers can be considered to comprise a team of sub-algorithms that collaborate to explore the space of design choices and associated hyper-parameters. By distributing the sub-algorithms across parallel computational resources, SMC samplers can improve time-efficiency. Since the sub-algorithms only need to avoid all failing at once, they can each be more adventurous in their exploration than the single MCMC algorithm: this can lead to energy-efficiency gains. Perhaps surprisingly, the potential for SMC samplers to automate design choices, while also exploring the associated hyper-parameter values, is largely unexplored. This PhD will investigate the significant potential to apply SMC samplers in this context.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金