Machine Learning through the Lenses of Optimal Transportation
Machine Learning through the Lenses of Optimal Transportation
批准号:
2744976
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --
中文摘要
本研究计划拟利用现代最优运输理论,探讨广义多层神经网络训练的一些分析和数值方面的问题。最佳大众交通在数学中是一个相对较新的领域,但在过去20年左右的时间里,它得到了巨大的发展势头,并在2018年获得了a . Figalli的菲尔兹奖。该理论在偏微分方程、概率论与统计、几何与数据科学等领域产生了巨大的影响和影响,在气象学、经济学、生物学和社会科学等领域都有重要的应用。应用于机器学习,在监督和无监督学习的框架下,最优大众交通已经取得了早期的成功。结果很快就找到了工业应用,例如基于变分自编码器和生成对抗网络的模型。训练神经网络的关键目标是识别一个函数,该函数将根据给定的训练集准确地评估未知数据。不幸的是,识别这个预测函数本身就是一个具有挑战性的优化问题,它只能通过数值方法近似地解决,并且有许多约束。过去,这是通过反向传播和随机梯度下降等技术实现的。然而,随着神经元数量的增加,这种方法不能很好地扩展,这使得它们在许多实际应用中不可行。由于最优传输的最新发展,现在有可能找到无限宽单层神经网络的预测函数。这是通过求解从所谓的瓦瑟斯坦梯度流中导出的偏微分方程来完成的,这是最佳运输的关键机制。然而,到目前为止,这个数学框架只适用于单层神经网络,这是Figalli等人最近承认的一个关键限制。为了解决这个开放的问题,我们寻求开发一个数学框架,作为这个项目的支柱,使用Wasserstein梯度流来寻找两层或多层神经网络上的预测函数。因此,最优运输的工具使我们能够逼近非常多神经元的问题,而不是通过更少的神经元,而是通过连续体(即无限多)的神经元,让人想起统计物理中的平均场模型。我们的第二个目标是量化这些近似有多好,不仅在通常的分析意义上获得通常很少实际使用的定性边界,而且在应用中观察到的情况下也有明确的边界。为此,我们计划使用平均领域游戏的技术(游戏邦注:这是另一位菲尔兹奖得主p - l。狮子和他的合作者)。数值计算将成为这个项目的重要组成部分,在第一个实例中测试我们的数学框架的发展(不是它是否正确,而是它如何直接适用于现实生活中的问题)。然而,现有的计算技术只能处理离散(如果有很多的话)神经元。在这个项目的过程中,作为我们的第三个目标,我们将需要开发数值模型,以及必要的计算技术,将离散神经元与(近似)连续神经元结合起来。这将从一个单层开始,随着我们获得经验而发展到多层。
英文摘要
This research project proposes to investigate some analytical and numerical aspects of the training of wide multi-layered neural networks using the modern theory of optimal transportation. Optimal mass transportation is a relatively new area in mathematics, but it has received a huge momentum in the past twenty years or so, culminating in the 2018 Fields medal of A. Figalli. This theory has had great impact and influence on fields as partial differential equations, probability theory and statistics, geometry and data science, with important applications in meteorology, economics, biology and social sciences, to list a few.Applied to machine learning, optimal mass transportation has had early successes, in the framework of supervised and unsupervised learning.The results have quickly found industrial applications, such as models based on variational auto-encoders and generative adversarial networks.The key objective in training a neural network is to identify a function that will accurately evaluate unseen data based on a given training set. Unfortunately, identifying this prediction function is itself a challenging optimisation problem, which can be solved only approximately and with many constraints through numerical methods. In the past, this has been achieved through techniques such as back propagation and stochastic gradient descent. However, such methods do not scale well as the number of neurons increases, making them infeasible for many practical applications.Thanks to recent developments in optimal transport, it is now possible to find prediction functions for infinitely wide single-layer neural networks.This is done by solving partial differential equations derived from so-called Wasserstein gradient flows, a key mechanism in optimal transport.Thus far, however, this mathematical framework is only applicable to single-layer neural networks, a key limitation acknowledged recently by Figalli et al.To address this open problem, we seek to develop, as the backbone of this project, a mathematical framework to use Wasserstein gradient flows to find prediction functions on neural networks with two or more layers.The tools of optimal transportation thus allow us to approximate the very-many-neuron problem, not by fewer neurons, but by a continuum (i.e. infinitely many) of them, reminiscent to the mean-field models in statistical physics. Our second objective is to quantify how good these approximations are, not only in the usual analytical sense of obtaining qualitative bounds that are often of little practical use, but also sharp bounds that align well with what is observed in applications. For this, we plan to use techniques from mean-field games (a new area initiated by another Fields medallist, P.-L. Lions, and his collaborators).Numerical computations will form an important part of this project, in the first instance to test our mathematical framework as it develops (not whether it is correct, but how directly applicable it is to real-life problems). Existing computational techniques, however, could only handle discrete (if many) neurons.Over the course of this project, as our third objective, we will need to develop numerical models, along with the requisite computational techniques, that couple discrete neurons with (an approximation of) a continuum of neurons. This will start with a single layer, progressing to multiple layers as we gain experience.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: