课题基金 / 基金详情

Collaborative Research: CIF: Medium: MoDL:Toward a Mathematical Foundation of Deep Reinforcement Learning

Collaborative Research: CIF: Medium: MoDL:Toward a Mathematical Foundation of Deep Reinforcement Learning
合作研究:CIF:媒介:MoDL:迈向深度强化学习的数学基础
批准号:
2212261
负责人:
Simon Du
金额:
$60.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-10-01 至 2026-09-30

项目摘要

项目成果

Simon Du的其他基金

相似基金

相关文献

中文摘要
翻译
深度强化学习(DRL)使用神经网络来解决顺序决策问题,在机器人、游戏、医疗保健和交通系统等现实世界的应用中取得了突破。然而,目前关于强化学习的理论工作仅限于状态数较少的问题;由于这些结果不包括神经网络,它们不能很好地解释DRL的经验成功。这个项目试图通过建立DRL的数学基础来弥合这一差距,该基础利用了近似理论、控制理论和优化理论的思想。这将使DRL的计算和统计复杂性得到系统的表征,并将有助于设计更有效和可靠的经验方法。该项目纳入了教育和外展计划。具体地说,调查人员将指导研究生和本科生(其中一些人是通过华盛顿大学针对代表性不足群体的STAR项目),开发新的课程和专著,组织研究研讨会,并为高中数据科学和人工智能课程开发课程材料。这个项目有三个主要组成部分。第一个重点确定了针对不同强化学习问题实例的策略可以实现哪些类型的保证。具体地说,这需要研究日益结构化的问题实例如何为策略提供更强有力的保证;这将通过使用并进一步开发非凸优化工具来实现,以描述实现奖励函数的固定点、局部极大值和全局极大值的策略。第二个推力从逼近理论和容量控制的角度来研究如何逐步增加神经网络的复杂性,以最终找到允许样本效率算法的最复杂的神经网络子族。第三个推力建立在前两个推力中获得的知识基础上,致力于设计计算效率高的算法;这将通过利用最优化理论的工具和与控制理论的联系来完成。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep Reinforcement Learning (DRL), which uses neural networks to solve sequential decision-making problems, has made breakthroughs in real-world applications, such as robotics, gaming, healthcare, and transportation systems. However, current theoretical work on reinforcement learning is restricted to problems with a small number of states; as these results do not cover neural networks, they cannot be used to satisfactorily explain the empirical successes of DRL. This project seeks to bridge this gap by building a mathematical foundation for DRL that leverages ideas from approximation theory, control theory, and optimization theory. This will allow the computational and statistical complexity of DRL to be systematically characterized, and will help with designing more efficient and reliable empirical methods. Education and outreach plans are integrated into this project. Specifically, the investigators will mentor graduate and undergraduate students (some through the STARS program for underrepresented groups at the University of washington), develop new courses and monographs, organize research workshops, and develop course materials for a high school data science and artificial intelligence curriculum. This project has three major components. The first thrust identifies which types of guarantees are achievable by policies for different reinforcement learning problem instances. Concretely, this requires investigating how increasingly structured problem instances enable stronger guarantees for policies; this will be done by using, and further developing, tools from non-convex optimization to describe policies that achieve stationary points, local maxima, and global maxima of the reward function. The second thrust takes the perspective of approximation theory and capacity control to investigate how the neural network complexity can be gradually increased to eventually find the most complex sub-family of neural networks that permit sample-efficient algorithms. The third thrust builds upon the knowledge gained in the first two thrusts, and is devoted to the design of computationally efficient algorithms; this will be done by leveraging tools from optimization theory and by making connections with control theory.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
DOI: --
发表时间: 2021-02
期刊:
影响因子: --
作者: [Zhihan Xiong;Ruoqi Shen;Qiwen Cui;Maryam Fazel;S. Du]
通讯作者: Zhihan Xiong;Ruoqi Shen;Qiwen Cui;Maryam Fazel;S. Du
DOI: 10.48550/arxiv.2206.01880
发表时间: 2022-06
期刊: ArXiv
影响因子: --
作者: [Qiwen Cui;Zhihan Xiong;Maryam Fazel;S. Du]
通讯作者: Qiwen Cui;Zhihan Xiong;Maryam Fazel;S. Du
On Controller Reduction in Linear Quadratic Gaussian Control with Performance Bounds
关于具有性能界限的线性二次高斯控制中的控制器简化
DOI: --
发表时间: 2023
期刊: Proceedings of Machine Learning Research
影响因子: --
作者: [Ren, Zhaolin, Zheng, Yang, Fazel, Maryam, Li, Na]
通讯作者: Li, Na
Toward a Theoretical Foundation of Policy Optimization for Learning Control Policies
为学习控制策略奠定策略优化的理论基础
DOI: 10.1146/annurev-control-042920-020021
发表时间: 2023
期刊: and Autonomous Systems
影响因子: --
作者: [Hu, Bin, Zhang, Kaiqing, Li, Na, Mesbahi, Mehran, Fazel, Maryam, Başar, Tamer]
通讯作者: Başar, Tamer
6
    CAREER: Toward a Foundation of Over-Parameterization
    • 批准号:
      2143493
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $57.0万
    • 财政年份:
      2022
    • 负责人:
      Simon Du
    • 依托单位:
    Collaborative Research: SCALE MoDL: Adaptivity of Deep Neural Networks
    • 批准号:
      2134106
    • 项目类别:
      Continuing Grant
    • 资助金额:
      $30.0万
    • 财政年份:
      2021
    • 负责人:
      Simon Du
    • 依托单位:
    IIS:RI Theoretical Foundations of Reinforcement Learning: From Tabula Rasa to Function Approximation
    • 批准号:
      2110170
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2021
    • 负责人:
      Simon Du
    • 依托单位:
    国内基金
    海外基金
    Research on Quantum Field Theory without a Lagrangian Description
    • 批准号:
      24ZR1403900
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
      SATOSHI NAWATA
    • 依托单位:
    Cell Research
    Cell Research
    Cell Research (细胞研究)