CAREER: Symbolic Learning with Neural Language Models
CAREER: Symbolic Learning with Neural Language Models
批准号:
2338833
负责人:
Kevin Ellis
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-05-01 至 2029-04-30
中文摘要
今天的人工智能(AI)系统在从大量数据中学习统计知识方面非常有效。然而,它们在学习我们人类可能用语言相互交流的知识形式方面效率较低,比如游戏规则、食物配方或填写税款的过程。这些形式的知识是符号的,这意味着它们可以用语言中的句子来表示,或者也可以用计算机代码来表示。在这个项目中,研究者将开发新的人工智能方法来学习表示为计算机代码的符号知识,结合统计学、大型语言模型(如ChatGPT)和程序合成(如何自动生成计算机软件)的思想。这项研究的科学影响将是人工智能系统从更少的例子中学习更抽象的知识形式,这对人类来说更容易理解,因为这些系统将用我们可以理解的语言描述它们所知道的东西。这项研究还将涉及学生研究人员,特别是那些来自代表性不足群体的研究人员。它还将为新的研究生和本科课程提供信息,包括新的康奈尔大学本科人工智能课程,每学期为大约150名学生提供服务。更详细地说,这项工作解决了学习符号知识的问题。符号表示已经成为自动化规划、证明助手和其他重要应用程序的基石,但与我们手动编码这些知识的能力相比,学习符号知识的能力还不太成熟。这项工作是围绕这样的观察组织的:像Python这样的通用编程语言在表示某些类型的符号知识方面非常有效,而且预训练的神经语言模型也擅长生成这样的代码。基于这些观察,该项目采用了一种框架,该框架结合了符号知识、用于不确定性估计的贝叶斯学习、程序合成和用于代码生成和有效概率推理的神经语言模型。建议的工作最终将有利于计划和基于模型的顺序决策,帮助我们更好地理解人类的思维和学习计算术语,并采取进一步自动化软件工程的步骤。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Artificial intelligence (AI) systems today are very effective at learning statistical knowledge from large amounts of data. However, they are less effective at learning the forms of knowledge that we as humans might communicate in language to each other, such as the rules of a game, a food recipe, or the process of filling out your taxes. These forms of knowledge are symbolic, meaning that they can be represented as sentences in language, or, alternatively, as computer code. In this project the investigator will develop new AI methods for learning symbolic knowledge represented as computer code, combining ideas from statistics, large language models such as ChatGPT, and program synthesis (how to automatically generate computer software). The scientific impact of this research will be AI systems that learn more abstract forms of knowledge, from fewer examples, and which are more understandable to humans, because the systems will describe what they know in languages we can understand. This research will also involve student researchers, especially those from underrepresented groups. It will also inform new graduate and undergraduate classes, including the new Cornell undergraduate AI class, which serves around 150 students each semester. In more detail, this work addresses the problem of learning symbolic knowledge. Symbolic representations already form the cornerstone of automated planning, proof assistants, and other important applications, but the ability to learn symbolic knowledge is less mature compared to our ability to manually encode such knowledge. The work is organized around the observation that general-purpose programming languages like Python are very effective at representing certain kinds of symbolic knowledge, and also that pretrained neural language models are adept at generating such code. Based on these observations, the project adopts a framing that combines symbolic knowledge, Bayesian learning for uncertainty estimation, program synthesis, and neural language models for code generation and efficient probabilistic inference. The proposed work could ultimately benefit planning and model-based sequential decision-making, help us better understand human thinking and learning in computational terms, and take steps toward further automating software engineering.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SHF: Small: Synthesizing Mixed Discrete/Continuous Programs with the Neurosymbolic Librarian
-
批准号:2310350
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2023
-
负责人:Kevin Ellis
-
依托单位:
海外基金