课题基金 / 基金详情

Regularized divergences and their gradient flows, generative modeling and structure-preserving learning.

Regularized divergences and their gradient flows, generative modeling and structure-preserving learning.
正则化散度及其梯度流、生成建模和结构保持学习。
批准号:
2307115
负责人:
Luc Rey-Bellet
金额:
$30.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-01 至 2026-07-31

项目摘要

项目成果

Luc Rey-Bellet的其他基金

相似基金

相关文献

中文摘要
翻译
生成性建模算法是人工智能最近和正在取得的许多进展的基础,无论是在流行的图像和文本生成工具方面,还是在材料设计、医学成像、药物发现和宇宙学等科学应用方面都是如此。这些算法的目标是从数据和可用的知识开始学习和构建模型,然后部署学习的模型来生成新的预测。这些预测是以新数据的形式进行的,例如新的图像、新的文本甚至是用于药物设计的新的候选分子。该项目涉及开发新的数学工具,从信息论、深度学习、微分方程和概率论到设计、改进、解释并最终信任这些学习算法。研究人员还将应用这些新算法来合并来自不同癌症研究的数据集,以解决通过整合来自相同疾病的数据来改善数据分析的迫切需求,这些数据是使用不同的研究、技术和患者组获得的。提出的研究的主要目标是在数据稀缺或获取成本高昂的情况下开发新的可靠的机器学习算法。作为研究项目的一部分,研究生将接受这一领域的培训。概率差异和度量是数学对象,旨在衡量不同概率模型之间或模型与数据之间的差异,特别擅长于非常高维的环境。需要仔细设计分歧,以构建最能描述可用数据的模型。这个项目将结合最优传输、信息论、偏微分方程和深度学习的工具来开发Lipschitz正则化发散,它介于Wasserstein度量和信息论发散(例如Kullback-Leibler发散)之间,并提供灵活的损失函数族来比较非绝对连续的概率度量。在机器学习应用中,人们经常需要建立算法来对奇异的目标分布进行建模,要么是由于其固有的性质,例如集中在低维结构上的概率,要么是因为它们通常只通过数据知道。这些新的分歧将与深度学习相结合,以在概率空间中建立能够将任何初始分布传输到目标数据集的梯度流。这些新方法还将适用于结构保持学习,出现在从医学到新分子设计的各种应用中,在这些应用中,数据可以显示对称性或物理约束。这种额外的知识将被考虑在概率发散中,以有效的方式建立保持结构的生成算法。每当数据稀缺和/或获取成本高昂时,就会研究和量化结构在生成算法中的重要作用。这项研究的示范领域之一是生物信息学,在那里,可用的真实数据集,即使涉及相同的疾病,由于预算限制或患者的有限可获得性,样本量较低,例如在罕见疾病的情况下。这一奖项反映了NSF的法定使命,并通过使用基金会的智力价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Generative modeling algorithms underlie many recent and ongoing advances in artificial intelligence, both in popular image and text generation tools as well as scientific applications such as materials design, medical imaging, drug discovery, and cosmology, to name a few. The goal of such algorithms is to learn and construct a model starting from data and available knowledge, and then deploy the learned model to generate new predictions. These predictions are in the form of new data such as new images, new text or even new candidate molecules for drug design. This project involves the development of new mathematical tools from information theory, deep learning, differential equations and probability theory to design, improve, explain and ultimately trust such learning algorithms. The investigators will also apply these new algorithms to merge data sets from different cancer studies, addressing a critical need to improve data analysis by integrating data from the same disease but which are obtained using different studies, technologies, and patient groups. The primary goal of the proposed research is to develop new reliable machine learning algorithms when data is scarce or expensive to obtain. Graduate students will be trained in this field as part of this research project.Probability divergences and metrics are mathematical objects designed to measure discrepancies between different probabilistic models or between models and data and are especially adept in very high-dimensional settings. Divergences need to be carefully designed to construct models which best describe the available data. This project will combine tools from optimal transport, information theory, partial differential equations and deep learning to develop Lipschitz regularized divergences which interpolate between Wasserstein metrics and information-theoretic divergences (e.g. the Kullback-Leibler divergence) and which provide flexible families of loss functions to compare non-absolutely continuous probability measures. In machine learning applications one often needs to build algorithms to model target distributions which are singular, either by their intrinsic nature such as probabilities concentrated on low dimensional structures and/or because they are often only known through data. These new divergences will be combined with deep learning to build gradient flows in a probability space which are capable of transporting any initial distribution to a target data set. These new methods will also be adapted for structure-preserving learning, arising in applications ranging from medicine to the design of new molecules, where data can exhibit symmetries or physical constraints. This additional knowledge will be taken into account in the probability divergence to build structure-preserving generative algorithms in an efficient way. The essential role of structure in generative algorithms will be studied and quantified whenever data is scarce and/or expensive to obtain. One of the demonstration areas of this research is in bioinformatics, where available real datasets, even when they involve the same disease, have low sample size due to budgetary constraints or limited availability of patients e.g., in the case of rare diseases.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1137/21m1434453
发表时间: 2021-07
期刊: ArXiv
影响因子: --
作者: [P. Birmpa;Jinchao Feng;M. Katsoulakis;Luc Rey-Bellet]
通讯作者: P. Birmpa;Jinchao Feng;M. Katsoulakis;Luc Rey-Bellet
Robust Uncertainty Quantification and Statistical Learning for Heavy Tails and Rare Events
  • 批准号:
    2008970
  • 项目类别:
    Standard Grant
  • 资助金额:
    $37.0万
  • 财政年份:
    2020
  • 负责人:
    Luc Rey-Bellet
  • 依托单位:
Mathematical and Computational Methods for Non-Equilbrium Systems
  • 批准号:
    1515712
  • 项目类别:
    Standard Grant
  • 资助金额:
    $27.97万
  • 财政年份:
    2015
  • 负责人:
    Luc Rey-Bellet
  • 依托单位:
Game Theory and Statistical Mechanics.
  • 批准号:
    1109316
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.5万
  • 财政年份:
    2011
  • 负责人:
    Luc Rey-Bellet
  • 依托单位:
AMC-SS: Mathematical and Computational in Nonequilibrium Statistical Mechanics.
  • 批准号:
    0605058
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $0.0万
  • 财政年份:
    2006
  • 负责人:
    Luc Rey-Bellet
  • 依托单位:
海外基金