Optimization methods for deep learning: training and testing in machine learning
Optimization methods for deep learning: training and testing in machine learning
批准号:
2282418
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --
中文摘要
这个项目属于EPSRC数值分析研究领域。这个项目是与工业合作伙伴NAG一起进行的。优化问题形成了机器学习和统计方法的建模和数值核心,例如在有监督的机器学习分类器的培训中。在这样的应用中,大量的数据(特征向量)是可用的,这些数据已经被分类(即标记)。然后选择一个参数化的分类器,并在该数据上进行训练,即计算参数的值,使得在给定特征处的分类器的输出以某种最优的方式匹配它们的标签。然后使用随后的分类器进行测试,以便对看不见的数据进行标记/分类。训练问题被表示为最小化例如分类器在测试集上犯下的平均错误量的优化问题。使用了优化问题的各种公式,最常见的考虑了一些连续损失函数来测量每个数据点的误差,以及损失项的(确定性)有限和或(概率)期望作为总误差;后者可以是凸的(例如在二进制分类的情况下),但由于深度学习应用的流行,它可能越来越多地改为非凸的。随之而来的优化规模通常是巨大的,在目标函数和中有数百万个参数和项。这使得计算单个函数或梯度值的成本高得令人望而却步,并导致需要有效地利用问题结构的不精确的优化算法。ML应用中的实用选择方法是(批量)随机梯度法(Robbins-Munro,1950),它只计算少量随机选择的损失项的梯度,并通过预先定义的步长控制方差和收敛。这一领域的一个重大挑战是如何利用不精确的二阶导数信息来增强随机梯度方法,从而获得更有效的方法,特别是在深度学习的非凸情况下,无论是在获得更高的精度方面,还是在对病态的鲁棒性方面。在这个项目中,我们将研究在ML优化问题的有限和结构中逼近二阶信息的方法,从次抽样二阶导数到利用批次梯度的差异来逼近它们,例如分块(随机)拟牛顿方法和高斯-牛顿方法。代替通常的预定义步长,我们将考虑经典优化中更复杂的步长技术的影响,这些技术适用于局部优化环境,如可变/自适应信赖域半径、正则化和线搜索。我们还将考虑包含廉价二阶导数信息的预条件技术,以帮助一阶方法的性能。我们还将研究这些方法的并行和分散实现,这对高阶技术来说是一个挑战。潜在的结果:(I)最新的深度学习和机器学习的优化公式和方法。(Ii)使用不精确(确定性/随机)问题信息的新的优化方法。(Iii)深度神经网络训练和测试方法的评估。
英文摘要
This project falls within the EPSRC numerical analysis research areas.This project is undertaken with Industrial Partner NAG.Optimization problems form the modelling and numerical core of machine learning and statistical methodologies, such as in the training of supervised machine learning classifiers. In such applications, a large amount of data (feature vectors) is available that has been already classified (namely labelled). Then a parameterised classifier is selected and trained on this data, namely, values of the parameters are calculated so that the output of the classifier at the given features matches their labels in some optimal way. The ensuing classifier is then used for testing, in order to label/classify unseen data. The training problem is formulated as an optimization problem that minimises, for example, the average amount of errors that the classifier makes on the test set. Various formulations of the optimization problem are used, most commonly considering some continuous loss function to measure the error at each data point and either (deterministic) finite sum or (probabilistic) expectation of the loss terms as the total error; the latter may be convex (such as in the case of binary classification) but, more and more nowadays, it may instead be nonconvex due to the prevalence of deep learning applications. The scale of the ensuing optimization is commonly huge, with millions of parameters and terms in the objective sum of functions. This makes the calculation of a single function or gradient value prohibitively expensive and leads to the need for inexact optimization algorithms that effectively exploit problem structure. The practical method of choice in ML applications is the (batch) stochastic gradient method (Robbins-Munro, 1950) that computes only the gradients of a small, randomly chosen number of the loss terms and controls both variance and convergence by means of a predefined stepsize. A grand challenge in this area is how to augment stochastic gradient methods with inexact second order derivative information, so as to obtain more efficient methods especially in the nonconvex case of deep learning, both in terms of achieving higher accuracy but also robustness to ill-conditioning. In this project, we will investigate ways to approximate second-order information in the finite-sum structure of ML optimization problems, from subsampling second-order derivatives to approximating them by differences in batch gradients, such as in block (stochastic) quasi-Newton approaches and Gauss-Newton methods. In place of the usual predefined stepsize, we will consider the impact of more sophisticated stepsize techniques from classical optimization that are adaptive to local optimization landscapes such as variable/adaptive trust-region radius, regularization and linesearch. We will also consider preconditioning techniques that contain inexpensive second-derivative information so as to help the performance of first order methods. We will also investigate parallel and decentralised implementations of these methods, which is a challenge especially for higher-order techniques.Potential outcomes:(i) State-of-the-art deep learning and optimization formulations and methods for machine learning.(ii) Novel optimization methods that use inexact (deterministic/stochastic) problem information.(iii) Evaluation of methods for deep neural net training and testing.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
复杂图像处理中的自由非连续问题及其水平集方法研究
-
批准号:60872130
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2008
-
负责人:刘国才
-
依托单位:
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: