Collaborative Research: RI: Medium: Superhuman Imitation Learning from Heterogeneous Demonstrations
Collaborative Research: RI: Medium: Superhuman Imitation Learning from Heterogeneous Demonstrations
批准号:
2312956
负责人:
Sanjiban Choudhury
金额:
$39.94万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2026-06-30
中文摘要
从示范行为(即模仿)中学习是在动物和人类中进行知识转移的一种有效手段。现有的人工智能(AI)系统的模仿学习方法通常假设模仿者的能力与演示者的能力匹配。当模仿者的能力远远超过示威者时,这可能会导致不受欢迎的行为。这个项目通过寻求使人工智能系统明确地比人类演示者更好,为在某些方面比(人类)演示者更有能力的人工智能系统重新制定了模仿学习。该项目将培训研究生和本科生开发人工智能系统,使其在广泛的高影响未来应用中更好地符合安全和实用要求。该项目使用最大限度优化控制/决策政策的引导(深度)强化学习来实现其重新制定的模仿学习目标。它侧重于从不同的演示和任务中学习,这些演示和任务在质量、难度和结构上存在差异。最初,假设可以使用多个指标来评估和比较不同的行为。在项目的稍后部分,将使用深度表示学习方法从演示和补充注释中学习这些指标。该项目方法产生的政策将在一系列不同的应用程序上进行评估:开源模拟器(例如,Atari游戏)、机器人平台的操作和移动任务,以及癌症治疗决策。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Learning from demonstrated behavior (i.e., imitation) is an effective means of knowledge transfer in animals and humans. Existing imitation learning methods for artificial intelligence (AI) systems typically assume the capabilities of the imitator match those of the demonstrator. This can lead to undesirable behavior when the imitator’s capabilities significantly exceed those of the demonstrator. This project reformulates imitation learning for AI systems that are more capable than (human) demonstrators in some aspects by seeking to make the AI system unambiguously better than human demonstrators. The project will train graduate students and undergraduates to develop artificial intelligence systems that are better aligned with safety and utility requirements in a broad range of highly impactful future applications.The project approaches its reformulated imitation learning objective using a maximum margin optimization for guiding (deep) reinforcement learning of control/decision policies. It focuses on learning from heterogeneous demonstrations and tasks that differ in quality, difficulty, and structure. Initially, multiple metrics for assessing and comparing different behaviors are assumed to be available. Later in the project, these metrics will be learned from demonstrations and supplemental annotations using deep representation learning methods. The policies produced by the approach of this project will be evaluated on a diverse set of applications: open source simulators (e.g., Atari games), manipulation and mobility tasks for robotics platforms, and cancer treatment decisions.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: Inverse Task Planning from Few-Shot Vision Language Demonstrations
-
批准号:2327973
-
项目类别:Standard Grant
-
资助金额:$50.73万
-
财政年份:2024
-
负责人:Sanjiban Choudhury
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: