课题基金 / 基金详情

Optimisation Methods for Optimal Transport

Optimisation Methods for Optimal Transport
最佳运输的优化方法
批准号:
2740715
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
最优运输是一个优雅的数学领域,它涉及到在特定的成本函数选择下,以最有效的方式将质量从一个概率分布转移到另一个概率分布的研究。它为底层空间上的成本函数提供了一种原则性机制,以推导出该空间上概率分布之间的距离的度量。这样的问题经常出现在机器学习系统中,要么是作为损失函数,要么是为了寻找分布之间的最优映射。虽然最优运输的数学理论在过去几十年中得到了很好的发展,但由于最近的计算进步,使得最优映射的计算变得容易处理,数据驱动的应用程序变得越来越重要。最优传输在机器学习系统中的最新应用包括图像生成、对齐单细胞数据和自然语言处理的方法,还有许多令人兴奋的应用有待探索。最优运输问题的解在很大程度上依赖于基础成本函数的选择。计算最优运输的大多数应用只考虑使用标准的二次成本,这通常是一个任意的选择,对于手头的问题可能不是一个好的成本函数。在这个项目中,我们的目标是利用经典的最优传输理论和最近的计算进展来设计方法,这些方法可以原则性地从观测数据中学习改进的成本函数,从而在后续的下游任务中提高性能。我们将使用统计学习理论的元素来提供收敛保证,确保我们的方法既可靠又高效。最优传输的一个应用是对单细胞组学数据的分析,单细胞组学数据包括对种群中单个细胞的测量,代价是在此过程中破坏细胞。因此,数据仅记录为接近整个人口的未标记快照。可以使用最佳传输方法来对齐观察到的分布,从而可以推断群体中单个细胞的轨迹。考虑到细胞测量是根据特定的矢量嵌入来记录的,这种表示的成本函数的好的选择是不清楚的,因此从现有数据学习成本函数的能力可以使性能得到改善。当底层空间的成本函数选择不清楚时,学习适合手头问题的成本函数的能力可能会在许多其他最优运输的应用和各种不同的数据结构中受益。由于计算最优运输方法在许多机器学习系统中发挥着重要作用,该建议属于EPSRC的《人工智能、数字化和数据:驱动价值和安全》的研究重点。
英文摘要
Optimal Transport is an elegant area of mathematics relating to the study of transporting the mass from one probability distribution to another in the most 'efficient' manner, with respect to a particular choice of cost function. It provides a principled mechanism for a cost function on the underlying space to induce a measure of distance between probability distributions on this space. Such problems occur frequently in machine learning systems, either as a loss function or to find an optimal mapping between distributions. While the mathematical theory of optimal transport has been well-developed over the past few decades, data-driven applications have become increasingly relevant due to recent computational advances that have allowed for tractable calculation of optimal mappings. Recent applications of optimal transport in machine learning systems have include methods for image generation, aligning single-cell data, and natural language processing, with many exciting applications yet to be explored. The solution to the optimal transport problem is highly dependent on the choice of underlying cost function. The majority of applications of computational optimal transport only consider using the standard quadratic cost, which is often an arbitrary choice and may not be a good cost function for the problem at hand. In this project, we aim to leverage both classical optimal transport theory and recent computational advances to design methods that can learn improved cost functions from observed data in a principled manner, which can thus lead to improved performance in subsequent downstream tasks. We will use elements of statistical learning theory to provide convergence guarantees, ensuring that our methods are both reliable and efficient.An application of optimal transport that could benefit from such methods is the analysis of single-cell omics data, which consists of measurements taken of individual cells in a population at a cost of destroying the cell in the process. The data is therefore recorded only as unlabelled snapshots that approximate the entire population. Optimal transport methods can be used to align the observed distributions, allowing the trajectories of individual cells in the population to be inferred. Given that the cell measurements are recorded according to a particular vector embedding, a good choice of cost function for this representation is not clear, so the ability to learn a cost function from existing data could enable improved performance. The ability to learn cost functions that are adapted for the problem at hand could be beneficial in many other applications of optimal transport and for a variety of different data structures, whenever the choice of cost function for the underlying space is unclear.As computational optimal transport methods play an important role in many machine learning systems, this proposal falls within the EPSRC's 'AI, Digitalisation and Data: Driving Value and Security' research priority.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Computational Methods for Analyzing Toponome Data