课题基金 / 基金详情

CAREER: Dual Reinforcement Learning: A Unifying Framework with Guarantees

CAREER: Dual Reinforcement Learning: A Unifying Framework with Guarantees
职业:双重强化学习:有保证的统一框架
批准号:
2340651
负责人:
Amy Zhang
金额:
$59.98万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-09-01 至 2029-08-31

项目摘要

项目成果

Amy Zhang的其他基金

相似基金

相关文献

中文摘要
翻译
强化学习(RL)有望自动化和改进许多需要顺序决策以优化某些长期目标的现实世界流程,例如自动驾驶汽车,工业自动化,推荐系统以及最近的自然语言处理。在过去的几年里,深度强化学习领域取得了令人兴奋的进展,强化学习代理在广泛的问题领域中表现出卓越的性能。然而,要实现这一进展,必须能够访问快速模拟器和数千万或数亿个数据点,这些数据点被收集、训练,然后被丢弃。非策略方法是一种替代方法,它提供了更高的数据效率,因为它们不仅限于在策略数据上进行训练,甚至可以用于在现有的离线数据上进行训练。这表明,要真正释放强化学习的潜力,我们必须开发有原则的非策略算法。这个项目的重点是通过研究一个框架来推进强化学习,该框架旨在提供一个统一的,原则性的目标,适用于标准和离线强化学习设置,并使我们能够有效地解决大规模,现实世界,顺序决策问题。在这个项目中,PI将研究这个目标的双重表述,这就产生了一个原则性的政策外目标,回避了更常用的原始公式中存在的问题。这一目标将导致算法特别适合于大的状态-动作空间,长的视野,和稀疏的奖励在现实世界中遇到的问题。PI将探索现有和新的模仿学习和模仿学习方法与拟议框架之间的联系。PI将表明,模仿学习和强化学习方法在这一目标下是统一的,并为这类方法提供理论保证。最后,PI将扩展双重框架,以利用预训练和微调来提高样本效率。这包括探索将域外数据集和多种模式纳入自我监督预训练的方法,特别是与家用机器人应用相关的方法。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Reinforcement learning (RL) holds the promise to automate and improve many real-world processes that require sequential decision-making to optimize some long-term objective, such as self-driving cars, industry automation, recommendation systems, and more recently in natural language processing. There has been much exciting progress in the field of deep reinforcement learning in the past few years, with RL agents demonstrating remarkable performance across a wide range of problem domains. However, to achieve this progress, it is necessary to have access to a fast simulator and tens or hundreds of millions of data points that are collected, trained on, then thrown away. Off-policy methods are an alternative approach, which provide much more data efficiency because they are not restricted to only training on on-policy data and can even be used to train on existing offline data. This suggests that to truly unlock the potential of reinforcement learning, we must develop principled off-policy algorithms. This project is focused on advancing RL by looking at a framework that aims to provide a unified, principled objective that applies to both standard and off-line RL settings and will allow us to efficiently solve large-scale, real-world, sequential decision-making problems.In this project, the PI will examine the dual formulation of this objective, which gives rise to a principled off-policy objective that sidesteps issues present in the more commonly used primal formulation. This objective will lead to algorithms particularly suitable for large state-action spaces, long horizons, and sparse rewards encountered in real-world problems. The PI will explore connections between existing and new imitation learning and reinforcement-learning methods and the proposed framework. The PI will show that both imitation learning and reinforcement learning methods are unified under this objective and present theoretical guarantees for this class of methods. Finally, the PI will extend the dual framework to leverage pre-training and fine tuning for improved sample efficiency. This includes exploring methods for incorporating out-of-domain datasets and multiple modalities in self-supervised pre-training, especially relevant for applications in household robotics.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CAREER: Tools for User and Community-Led Social Media Curation
  • 批准号:
    2236618
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $61.28万
  • 财政年份:
    2023
  • 负责人:
    Amy Zhang
  • 依托单位:
Collaborative Research: DASS: Transitioning open-source software projects to accountable community governance
  • 批准号:
    2217653
  • 项目类别:
    Standard Grant
  • 资助金额:
    $15.0万
  • 财政年份:
    2022
  • 负责人:
    Amy Zhang
  • 依托单位:
Collaborative Research: SaTC: CORE: Large: Privacy-Preserving Abuse Prevention for Encrypted Communications Platforms
  • 批准号:
    2120497
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $37.14万
  • 财政年份:
    2021
  • 负责人:
    Amy Zhang
  • 依托单位:
国内基金
海外基金
基于双荧光结核菌和Dual RNA-seq技术的病原宿主免疫互作关键基因挖掘及机制研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    史琛彦
  • 依托单位:
Dual AGN 的系统搜寻及其性质研究
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    张杨威
  • 依托单位:
AKAP3通过其Dual和RI结构域整合多重信号通路调控精子活力和男性育性的机理研究
  • 批准号:
    82171602
  • 项目类别:
    面上项目
  • 资助金额:
    54万元
  • 批准年份:
    2021
  • 负责人:
    徐凯彪
  • 依托单位:
磷化双贱金属合金超薄膜dual-(Bimetallene-P)催化材料的超声脉冲界面构筑及其电解水性能研究
  • 批准号:
    --
  • 项目类别:
    面上项目
  • 资助金额:
    60万元
  • 批准年份:
    2021
  • 负责人:
    温鸣
  • 依托单位: