课题基金 / 基金详情

Collaborative Research: RI: Medium: Bootstrapping natural feedback for reinforcement learning

Collaborative Research: RI: Medium: Bootstrapping natural feedback for reinforcement learning
合作研究:RI:中:引导强化学习的自然反馈
批准号:
2212310
负责人:
Jacob Andreas
金额:
$120.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2025-08-31

项目摘要

项目成果

Jacob Andreas的其他基金

相似基金

相关文献

中文摘要
翻译
人工智能的许多现代应用-从工业自动化到内容推荐-依赖于机器学习算法,这些算法训练自动代理与其环境交互。但交互式学习的两种主要方法--强化学习和模仿--需要如此多的监督或培训时间,以至于将它们应用于大多数现实世界的问题成本高得令人望而却步。人类的学习没有受到这一缺陷的影响,这在很大程度上是因为人类不是通过奖励或示范来学习,而是通过与使用手势和语言等信号的熟练教师进行长期互动来学习。本项目将从个体智能体、人-智能体团队和多智能体群体的角度,为研究具有丰富反馈的交互式学习奠定基础。它将为自动代理的互动培训产生新的能力,扩大此类技术的有效性和可及性。对自然、交互反馈的支持也将提高这类系统的定制化能力,使用户在没有显著计算能力、数据注释资源或甚至编程能力的情况下,可以随时进行调整或再培训。该项目分为三个广泛的研究目标。首先,它将开发一个正式的反馈基础框架,使用简单的监督信号(在执行期间或执行后提供)来引导对更复杂的反馈类型的学习解释。其次,它将开发学习算法,以征求反馈。这些算法将把强化学习的单向过程转变为双向交互,使代理能够主动向主管查询有关环境的组成和因果结构的信息。第三,它将开发提供反馈的新机制和技术,通过软件工具帮助人类主管选择或生成信息量最大的反馈信号。每个目标下的研究将在模拟环境中进行,使用跨越导航、机器人操作和家具组装的复杂任务进行基准测试,并根据其对样本效率、端到端开发时间和可用性的好处进行评估。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Many modern applications of artificial intelligence---from industrial automation to content recommendation---depend on machine learning algorithms that train automated agents to interact with their environments. But the two main approaches to interactive learning, reinforcement learning and imitation, require so much supervision or training time that it is prohibitively expensive to apply them to most real-world problems. Human learning does not suffer from this shortcoming, in large part because humans learn not from rewards or demonstrations, but instead from extended interaction with skilled teachers who use signals like gesture and language. This project will lay a foundation for research on interactive learning with rich feedback, from the perspective of individual agents, human--agent teams, and multi-agent populations. It will yield new capabilities for interactive training of automated agents, expanding both the effectiveness and accessibility of such techniques. Support for natural, interactive feedback will also improve the customizability of such systems, making on-the-fly adaptation or retraining accessible to users without significant computing power, data annotation resources or even programming ability.The project is organized into three broad research objectives. First, it will develop a formal framework for grounding feedback, using simple supervisory signals (provided during or after execution) to bootstrap learned interpretation of more complex feedback types. Second, it will develop algorithms for learning to solicit feedback. These algorithms will turn the one-way process of reinforcement learning into a two-way interaction, enabling agents to proactively query supervisors for information about the compositional and causal structure of the environment. Third, it will develop new mechanisms and techniques for providing feedback, via software tools that assist human supervisors in selecting or generating maximally informative feedback signals. Research under each of these objectives will be carried out in simulated environments, benchmarked using complex tasks spanning navigation, robot manipulation, and furniture assembly, and evaluated in terms of its benefits to sample efficiency, end-to-end development time, and usability.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI: 10.48550/arxiv.2302.06692
发表时间: 2023-02
期刊:
影响因子: --
作者: [Yuqing Du;Olivia Watkins;Zihan Wang;Cédric Colas;Trevor Darrell;P. Abbeel;Abhishek Gupta;Jacob Andreas]
通讯作者: Yuqing Du;Olivia Watkins;Zihan Wang;Cédric Colas;Trevor Darrell;P. Abbeel;Abhishek Gupta;Jacob Andreas
CAREER: Learning Structured Models with Natural Language Supervision
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)