课题基金 / 基金详情

Collaborative Research: CISE: Large: Executing Natural Instructions in Realistic Uncertain Worlds

Collaborative Research: CISE: Large: Executing Natural Instructions in Realistic Uncertain Worlds
合作研究:CISE:大型:在现实的不确定世界中执行自然指令
批准号:
2321852
负责人:
Tucker Hermans
金额:
$93.75万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2028-09-30

项目摘要

项目成果

Tucker Hermans的其他基金

相似基金

相关文献

中文摘要
翻译
为了让机器人像人类助手一样流畅地操作,它们必须能够从人类那里接受自然语言指令,并在复杂、不确定的环境中采取行动来实现这些指令。虽然有些商用机器人在物理上能够执行各种各样的有用指令,但目前的人工智能(AI)框架无法完全理解并自主执行大多数指令。相反,目前用于机器人指令遵循的人工智能技术通常局限于高度受限、不现实的环境和不自然、脆弱和僵化的指令格式。该项目的首要目标是研究基本的人工智能原理,使机器人能够在现实的、不确定的世界中可靠地执行自然指令。该项目有可能大幅增加社会可用的体力劳动,而不增加人力劳动的数量,通过使典型的人类工人能够指挥半自动机器人完成大量平凡的任务。只有当机器在自然环境中容易被人类指导,而人类只需要很少的专业训练,这才有可能。考虑到美国建设物理基础设施的能力需要增加,这种设想的劳动力乘数是极其相关的。例如,这包括新建和升级公共基础设施,如桥梁和能源系统,以及高效地建造负担得起的住房。这些进步还将对其他经济领域产生更广泛的影响,如物流、医疗保健、家庭助理。该项目还将通过K-12倡议、本科生研究经验和招募代表性不足的研究生人才,为教育和推广做出贡献。该项目将设计和开发一个新的集成框架,用于嵌入人工智能代理,该框架由计算机视觉、语言理解、世界建模、规划和控制方面的协同进步组成。该框架将在逐步增加能力的阶段计划中发展,从逐步执行指令开始,并逐步执行一般类型的目标导向指令。该研究将使用商用机器人在物理逼真的模拟环境和现实世界环境中测试和演示该框架。此外,将在每个功能阶段进行用户研究,以将工作重点放在最终用户效用上。该框架的核心是一种新的时空场景知识结构,即多模态实体图(MEM),它基于视觉和语言进行更新,并用于规划和技能执行。该研究将研究3D视觉和语言理解方面的新思路,以便在现实输入的基础上持续维护MEM,以捕捉环境中的不确定性。该项目还将研究机器人技能的低水平全身控制的新方法,该方法受到最近语言建模成功的启发,有助于模块化技能学习和知识共享。最后,该项目将通过研究MEMs上学习动态模型的新思路来推进自动化规划能力,这些新思路被用于基于动态条件语言模型的高水平技能规划的新方法。重要的是,所有这些创新都将以同步的方式开发,以便对集成框架进行严格的测试和演示。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
For robots to fluidly operate as human assistants they must be able to take natural language instructions from humans and act to achieve those instructions in complex, uncertain environments. While there are commodity robots that are physically capable of carrying out a wide variety of useful instructions, current artificial intelligence (AI) frameworks are not able to fully understand and autonomously execute most of those instructions. Rather, current AI techniques for robot instruction following have generally been limited to highly constrained, unrealistic environments and instruction formats that are unnatural, brittle, and rigid. The overarching goal of the project is to study the fundamental AI principles that enable robots to reliably execute natural instructions in realistic, uncertain worlds. The project has the potential to dramatically increase the physical labor available to society, without increasing the amount of human labor, by enabling typical human workers to direct semi-automated robots for a multitude of mundane tasks. This will only be possible if the machines are easily instructable in natural environments by humans who only require minimal specialized training. This envisioned labor multiplier is extremely relevant given the need for increased capacity to build physical infrastructure in the US. This includes, for example, new and upgraded public infrastructure such as bridges and energy systems, as well as efficient construction of affordable housing. These same advances will also result in broader impacts to other parts of the economy, such as logistics, healthcare, household assistants. The project will also contribute to education and outreach through K-12 initiatives, undergraduate research experiences, and recruiting of underrepresented graduate student talent.The project will design and develop a novel integrated framework for embodied AI agents that is comprised of synergistic advances in computer vision, language understanding, world modeling, planning, and control. The framework will evolve over a staged plan of increasing capabilities, starting with step-by-step instruction execution and progressing to executing general types of goal-oriented instructions. The research will test and demonstrate the framework in both physically-realistic simulation environments and real-world environments using commodity robots. In addition, user studies will be conducted at each capability stage to focus the work toward end-user utility. Central to the framework is a new knowledge structure for spatio-temporal scenes, the multi-modal entity map (MEM), which is updated based on vision and language and used for both planning and skill execution. The research will study new ideas in 3D vision and language understanding for continually maintaining the MEM based on realistic inputs in a way that captures uncertainty in the environment. The project will also study a new approach to low-level full-body control for robot skills, inspired by recent successes in language modeling, that facilitates both modular skill learning and knowledge sharing. Finally, the project will advance automated planning capabilities by studying new ideas for learning dynamics models over the MEMs, which are used by a novel approach to high-level skill planning based on dynamics-conditioned language models. Importantly all of these innovations will be developed in a synchronized way to allow for rigorous testing and demonstration of the integrated framework.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: NRI: FND: Learning Graph Neural Networks for Multi-Object Manipulation
  • 批准号:
    2024778
  • 项目类别:
    Standard Grant
  • 资助金额:
    $34.43万
  • 财政年份:
    2020
  • 负责人:
    Tucker Hermans
  • 依托单位:
CAREER: Improving Multi-Fingered Manipulation by Unifying Learning and Planning
  • 批准号:
    1846341
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $53.27万
  • 财政年份:
    2019
  • 负责人:
    Tucker Hermans
  • 依托单位:
CRII: RI: Enabling Manipulation of Object Collections via Self-Supervised Robot Learning
  • 批准号:
    1657596
  • 项目类别:
    Standard Grant
  • 资助金额:
    $17.5万
  • 财政年份:
    2017
  • 负责人:
    Tucker Hermans
  • 依托单位:
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)