CAREER: Theoretical Foundations of Modern Machine Learning Paradigms: Generative and Out-of-Distribution
CAREER: Theoretical Foundations of Modern Machine Learning Paradigms: Generative and Out-of-Distribution
批准号:
2238523
负责人:
Andrej Risteski
金额:
$52.95万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-03-01 至 2028-02-29
中文摘要
在过去的几年里,大规模机器学习模型的性能有了显著的提高。像DALL-E这样经过大量文本和图像数据训练的模型,只要给出图像内容的口头描述(“提示”),就可以生成高质量的图像,即使提示与模型训练的内容相距甚远。像ChatGPT这样的模型,经过大量文本数据和与人类互动的训练,能够像人类对话者一样令人信服地表现出来,甚至可以解决简单的数学考试问题。尽管这些系统代表了一项令人印象深刻的工程壮举,但我们对训练它们的配方中哪些成分是重要的科学理解严重缺乏。因此,改进它们通常涉及大量的试验和错误,这也转化为大量的人力时间和计算资源。我们对它们的失效模式(甚至常常是如何评估它们!)的理解甚至更差。因此,目前不太可能在任何安全关键场景中部署它们。该项目的目标是为理解现代机器学习模型的故障模式建立科学和数学基础,特别是在它们所训练和部署的数据之间存在变化的情况下。这个项目将概述在建立理解和改进现代生成模型的数学和科学基础方面的挑战,以及形式化可以处理与训练分布有很大不同的分布的设置。研究者将建立有用的学习范式的形式化,使其易于处理的结构性假设,并为此类设置开发新的算法工具。该项目将有两个主要方面:(1)在概率生成模型的背景下,用于采样、推理和学习的算法工具,以及开发用于理解不同类型模型的失效模式的分析机制和规避和改进它们的算法解决方案。(2)涉及数据分布变化的学习设置,包括领域泛化(当呈现来自多个数据环境的数据时,期望学习者在新环境中表现良好),领域翻译(当模型呈现来自两个领域的数据并期望学习如何在它们之间“翻译”时)和持续学习(当来自不同环境的数据以在线方式呈现时)。学习者被期望在所有环境中都保持良好的表现)。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The performance of large-scale machine learning models in the last several years has improved dramatically. Models like DALL-E, trained on vast amounts of text and image data, can generate high-quality images given just a verbal description ("prompt") of the image's content, even for prompts that are very far from what the model was trained on. Models like ChatGPT, trained on vast amounts of text data and interactions with humans, are capable of convincingly behaving like a human interlocutor, even solving simple mathematics exam questions. Though these systems represent an impressive feat of engineering, our scientific understanding of what ingredients in the recipe to train them are important is severely lacking. Thus, improving them often involves extensive trial and error, which also translates to considerable amounts of human hours and computational resources. Our understanding of their failure modes (or often even how to evaluate them!) is even poorer. Thus, deploying them in any safety-critical scenario is, at present, unlikely. The goal of this project is to build scientific and mathematical foundations for understanding failure modes of modern machine learning models, especially in the presence of changes between the data they are trained on and deployed on. This project will outline the challenges in building mathematical and scientific footing for understanding and improving modern generative models, as well as formalize settings in which distributions substantially different from the training distribution can be handled. The investigator will establish useful formalizations of learning paradigms, structural assumptions that make them tractable, and develop new algorithmic tools for such settings. The project will have two major prongs: (1) Algorithmic tools for sampling, inference, and learning in the context of probabilistic generative models, as well as developing analytical machinery for understanding failure modes of different families of models and algorithmic solutions to circumvent and ameliorate them. (2) Learning settings involving shifts in the data distribution, including domain generalization (when data from several data environments is presented, and the learner is expected to perform well in a new environment), domain translation (when a model is presented data from two domains and is expected to learn how to "translate" between them) and continual learning (when data from different environments is presented in an online fashion, and the learner is expected to retain good performance in all environments).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Provable Benefits of Score Matching
分数匹配的可证明的好处
DOI:
--
发表时间:
2023
期刊:
2023
影响因子:
--
作者:
[Pabbaraju, C, Rohatgi, D, Sevekari, A, Lee, H, Moitra, A, Risteski, A.]
通讯作者:
Risteski, A.
DOI:
10.48550/arxiv.2210.00726
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
作者:
[Frederic Koehler;Alexander Heckett;Andrej Risteski]
通讯作者:
Frederic Koehler;Alexander Heckett;Andrej Risteski
DOI:
10.48550/arxiv.2312.01429
发表时间:
2023-12
期刊:
ArXiv
影响因子:
--
作者:
[Kaiyue Wen;Yuchen Li;Bing Liu;Andrej Risteski]
通讯作者:
Kaiyue Wen;Yuchen Li;Bing Liu;Andrej Risteski
DOI:
10.48550/arxiv.2310.03902
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
作者:
[O. Chehab;Aapo Hyvarinen;Andrej Risteski]
通讯作者:
O. Chehab;Aapo Hyvarinen;Andrej Risteski
海外基金