课题基金 / 基金详情

Collaborative Research: CISE: Large: Executing Natural Instructions in Realistic Uncertain Worlds

Collaborative Research: CISE: Large: Executing Natural Instructions in Realistic Uncertain Worlds
合作研究:CISE:大型:在现实的不确定世界中执行自然指令
批准号:
2321852
负责人:
Tucker Hermans
金额:
$93.75万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2028-09-30

项目摘要

项目成果

Tucker Hermans的其他基金

相似基金

相关文献

中文摘要
翻译
为了让机器人流畅地充当人类助手,它们必须能够接受人类的自然语言指令,并在复杂、不确定的环境中采取行动实现这些指令。虽然有一些商用机器人在物理上能够执行各种有用的指令,但目前的人工智能(AI)框架无法完全理解并自主执行其中的大多数指令。相反,目前用于机器人指令跟踪的人工智能技术通常仅限于高度受限、不现实的环境和不自然、脆弱和僵硬的指令格式。该项目的首要目标是研究基本的人工智能原理,使机器人能够在现实、不确定的世界中可靠地执行自然指令。该项目有可能在不增加人类劳动量的情况下,显著增加社会可用的体力劳动,使典型的人类工人能够指导半自动机器人执行大量平凡的任务。只有当机器在自然环境中很容易被只需要最少专门训练的人教授时,这才是可能的。考虑到美国需要增加产能来建设实体基础设施,这一设想的劳动力乘数极其相关。例如,这包括新的和升级的公共基础设施,如桥梁和能源系统,以及高效建造负担得起的住房。这些同样的进步还将对其他经济领域产生更广泛的影响,如物流、医疗保健、家庭助理。该项目还将通过K-12倡议、本科生研究经验和招聘未被充分代表的研究生人才来促进教育和推广。该项目将设计和开发一个新的集成框架,用于具体化人工智能代理,包括在计算机视觉、语言理解、世界建模、规划和控制方面的协同进展。该框架将在一个增加能力的分阶段计划中发展,从循序渐进的指令执行开始,到执行一般类型的面向目标的指令。这项研究将使用商用机器人在物理逼真的模拟环境和现实世界环境中测试和演示该框架。此外,将在每个能力阶段进行用户研究,以将工作重点放在最终用户效用上。该框架的核心是时空场景的一种新的知识结构,即多通道实体映射(MEM),它基于视觉和语言进行更新,用于规划和技能执行。这项研究将研究3D视觉和语言理解方面的新想法,以捕捉环境中的不确定性,从而继续保持基于现实输入的MEM。该项目还将研究一种新的低水平全身控制机器人技能的方法,该方法的灵感来自最近在语言建模方面的成功,该方法促进了模块化技能学习和知识共享。最后,该项目将通过研究在MEMS上学习动力学模型的新想法来提高自动化规划能力,MEMS被一种基于动态条件语言模型的高级技能规划的新方法所使用。重要的是,所有这些创新都将以同步的方式开发,以允许对综合框架进行严格的测试和演示。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
For robots to fluidly operate as human assistants they must be able to take natural language instructions from humans and act to achieve those instructions in complex, uncertain environments. While there are commodity robots that are physically capable of carrying out a wide variety of useful instructions, current artificial intelligence (AI) frameworks are not able to fully understand and autonomously execute most of those instructions. Rather, current AI techniques for robot instruction following have generally been limited to highly constrained, unrealistic environments and instruction formats that are unnatural, brittle, and rigid. The overarching goal of the project is to study the fundamental AI principles that enable robots to reliably execute natural instructions in realistic, uncertain worlds. The project has the potential to dramatically increase the physical labor available to society, without increasing the amount of human labor, by enabling typical human workers to direct semi-automated robots for a multitude of mundane tasks. This will only be possible if the machines are easily instructable in natural environments by humans who only require minimal specialized training. This envisioned labor multiplier is extremely relevant given the need for increased capacity to build physical infrastructure in the US. This includes, for example, new and upgraded public infrastructure such as bridges and energy systems, as well as efficient construction of affordable housing. These same advances will also result in broader impacts to other parts of the economy, such as logistics, healthcare, household assistants. The project will also contribute to education and outreach through K-12 initiatives, undergraduate research experiences, and recruiting of underrepresented graduate student talent.The project will design and develop a novel integrated framework for embodied AI agents that is comprised of synergistic advances in computer vision, language understanding, world modeling, planning, and control. The framework will evolve over a staged plan of increasing capabilities, starting with step-by-step instruction execution and progressing to executing general types of goal-oriented instructions. The research will test and demonstrate the framework in both physically-realistic simulation environments and real-world environments using commodity robots. In addition, user studies will be conducted at each capability stage to focus the work toward end-user utility. Central to the framework is a new knowledge structure for spatio-temporal scenes, the multi-modal entity map (MEM), which is updated based on vision and language and used for both planning and skill execution. The research will study new ideas in 3D vision and language understanding for continually maintaining the MEM based on realistic inputs in a way that captures uncertainty in the environment. The project will also study a new approach to low-level full-body control for robot skills, inspired by recent successes in language modeling, that facilitates both modular skill learning and knowledge sharing. Finally, the project will advance automated planning capabilities by studying new ideas for learning dynamics models over the MEMs, which are used by a novel approach to high-level skill planning based on dynamics-conditioned language models. Importantly all of these innovations will be developed in a synchronized way to allow for rigorous testing and demonstration of the integrated framework.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: NRI: FND: Learning Graph Neural Networks for Multi-Object Manipulation
  • 批准号:
    2024778
  • 项目类别:
    Standard Grant
  • 资助金额:
    $34.43万
  • 财政年份:
    2020
  • 负责人:
    Tucker Hermans
  • 依托单位:
CAREER: Improving Multi-Fingered Manipulation by Unifying Learning and Planning
  • 批准号:
    1846341
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $53.27万
  • 财政年份:
    2019
  • 负责人:
    Tucker Hermans
  • 依托单位:
CRII: RI: Enabling Manipulation of Object Collections via Self-Supervised Robot Learning
  • 批准号:
    1657596
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.5万
  • 财政年份:
    2017
  • 负责人:
    Tucker Hermans
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)