Optimization methods for deep learning: training and testing in machine learning
Optimization methods for deep learning: training and testing in machine learning
批准号:
2282418
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
该项目福尔斯属于EPSRC数值分析研究领域。该项目与工业合作伙伴NAG合作。优化问题形成了机器学习和统计方法的建模和数值核心,例如在监督机器学习分类器的训练中。在这样的应用中,大量的数据(特征向量)已经被分类(即标记)。然后选择参数化分类器并在此数据上进行训练,即计算参数的值,使得给定特征处的分类器的输出以某种最佳方式匹配其标签。随后的分类器用于测试,以标记/分类看不见的数据。训练问题被公式化为最小化的优化问题,例如,分类器在测试集上产生的平均错误量。使用优化问题的各种公式,最常见的是考虑一些连续损失函数来测量每个数据点的误差,(确定性)有限和或(概率)损失项的期望值作为总误差;后者可以是凸起的(例如在二元分类的情况下),但是,现在越来越多,由于深度学习应用的流行,它可能是非凸的。随之而来的优化的规模通常是巨大的,在目标函数的总和中有数百万个参数和项。这使得单个函数或梯度值的计算过于昂贵,并导致需要有效利用问题结构的不精确优化算法。ML应用中的实际选择方法是(批量)随机梯度方法(Robbins-Munro,1950),它只计算少量随机选择的损失项的梯度,并通过预定义的步长控制方差和收敛。这一领域的一个巨大挑战是如何用不精确的二阶导数信息来增强随机梯度方法,以便获得更有效的方法,特别是在深度学习的非凸情况下,无论是在实现更高的准确性方面,还是在对病态的鲁棒性方面。在这个项目中,我们将研究在ML优化问题的有限和结构中近似二阶信息的方法,从二阶导数的子采样到通过批量梯度的差异近似它们,例如块(随机)拟牛顿方法和高斯-牛顿方法。代替通常的预定义步长,我们将考虑来自经典优化的更复杂的步长技术的影响,这些技术适应于局部优化景观,例如可变/自适应信赖域半径,正则化和线性优化。我们还将考虑预处理技术,包含廉价的二阶导数信息,以帮助一阶方法的性能。我们还将研究这些方法的并行和分散实现,这对于高阶技术来说尤其是一个挑战。潜在成果:(i)最先进的深度学习和优化公式以及机器学习方法。(ii)使用不精确(确定性/随机性)问题信息的新优化方法。(iii)深度神经网络训练和测试方法的评估。
英文摘要
This project falls within the EPSRC numerical analysis research areas.This project is undertaken with Industrial Partner NAG.Optimization problems form the modelling and numerical core of machine learning and statistical methodologies, such as in the training of supervised machine learning classifiers. In such applications, a large amount of data (feature vectors) is available that has been already classified (namely labelled). Then a parameterised classifier is selected and trained on this data, namely, values of the parameters are calculated so that the output of the classifier at the given features matches their labels in some optimal way. The ensuing classifier is then used for testing, in order to label/classify unseen data. The training problem is formulated as an optimization problem that minimises, for example, the average amount of errors that the classifier makes on the test set. Various formulations of the optimization problem are used, most commonly considering some continuous loss function to measure the error at each data point and either (deterministic) finite sum or (probabilistic) expectation of the loss terms as the total error; the latter may be convex (such as in the case of binary classification) but, more and more nowadays, it may instead be nonconvex due to the prevalence of deep learning applications. The scale of the ensuing optimization is commonly huge, with millions of parameters and terms in the objective sum of functions. This makes the calculation of a single function or gradient value prohibitively expensive and leads to the need for inexact optimization algorithms that effectively exploit problem structure. The practical method of choice in ML applications is the (batch) stochastic gradient method (Robbins-Munro, 1950) that computes only the gradients of a small, randomly chosen number of the loss terms and controls both variance and convergence by means of a predefined stepsize. A grand challenge in this area is how to augment stochastic gradient methods with inexact second order derivative information, so as to obtain more efficient methods especially in the nonconvex case of deep learning, both in terms of achieving higher accuracy but also robustness to ill-conditioning. In this project, we will investigate ways to approximate second-order information in the finite-sum structure of ML optimization problems, from subsampling second-order derivatives to approximating them by differences in batch gradients, such as in block (stochastic) quasi-Newton approaches and Gauss-Newton methods. In place of the usual predefined stepsize, we will consider the impact of more sophisticated stepsize techniques from classical optimization that are adaptive to local optimization landscapes such as variable/adaptive trust-region radius, regularization and linesearch. We will also consider preconditioning techniques that contain inexpensive second-derivative information so as to help the performance of first order methods. We will also investigate parallel and decentralised implementations of these methods, which is a challenge especially for higher-order techniques.Potential outcomes:(i) State-of-the-art deep learning and optimization formulations and methods for machine learning.(ii) Novel optimization methods that use inexact (deterministic/stochastic) problem information.(iii) Evaluation of methods for deep neural net training and testing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
复杂图像处理中的自由非连续问题及其水平集方法研究
-
批准号:60872130
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2008
-
负责人:刘国才
-
依托单位:
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: