课题基金 / 基金详情

CAREER: Reinforcement Learning for Recursive Markov Decision Processes and Beyond

CAREER: Reinforcement Learning for Recursive Markov Decision Processes and Beyond
职业:递归马尔可夫决策过程及其他的强化学习
批准号:
2146563
负责人:
Ashutosh Trivedi
金额:
$59.66万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-05-01 至 2027-04-30

项目摘要

项目成果

Ashutosh Trivedi的其他基金

相似基金

相关文献

中文摘要
翻译
该奖项的全部或部分资金来自《2021年美国救援计划法案》(公法117-2)。强化学习(RL)是一种基于抽样的马尔可夫决策过程(MDP)优化方法,代理依靠奖励来发现最优解决方案。当与强大的近似方案(例如,深度神经网络)相结合时,RL在传统上被认为超出人工智能能力范围的高度复杂的任务中一直是有效的。然而,它对近似参数的敏感性使得RL很难使用(程序员需要大量的机器学习专业知识),也很难信任(手动近似可能会使保证无效)。这个项目的愿景是通过开发原则性的方法和强大的工具来提高基于RL的编程的规模的可用性和可信性,从而使RL民主化。除了这些研究目标,还努力在CS教育中整合基于RL的可计算性的基础,并探索基于RL的编程在CS教育中的作用。由于RL算法在有限的MDP上工作,但可伸缩性较差,因此需要在RL中进行近似。近似性会影响可用性和可信度。该方案确定了解决这两个问题的两个目标:1)发现超越有限MDP的收敛的RL;2)开发具有严格优化保证的基于抽象的RL方法。所提议的方法的成功将通过它们处理大规模系统的能力来评估。算法和数据集将作为开放源码软件分发。这项拟议的研究对三个学科做出了根本性的贡献:形式方法、机器学习和控制理论;同时,它通过使基于RL的编程更容易和更具包容性,向扩大计算的参与迈出了根本、具体的步骤。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This award is funded in whole or in part under the American Rescue Plan Act of 2021 (Public Law 117-2).Reinforcement Learning (RL) is a sampling-based approach to optimization of Markov decision processes (MDPs), where agents rely on rewards to discover optimal solutions. When combined with powerful approximation schemes (e.g., deep neural networks), RL has been effective in highly complex tasks traditionally considered beyond reach of Artificial Intelligence. However, its sensitivity to the approximation parameters makes RL difficult to use (significant Machine Learning expertise is demanded of the programmer) and difficult to trust (manual approximations can invalidate guarantees). The vision of this project is to democratize RL by developing principled methodologies and powerful tools to improve the usability and trustworthiness of RL-based programming at scale. These research objectives are complemented by efforts to integrate the foundations of RL-based computability in CS education and to explore the role of RL-based programming in CS education.Approximation in RL is needed because RL algorithms with guaranteed convergence work on finite MDPs, and yet scale poorly. Approximation affects usability and trustworthiness. This proposal identifies two goals addressing both concerns: 1) to discover convergent RL beyond finite MDPs and 2) to develop abstraction-based approaches for RL with rigorous optimization guarantees. The success of the proposed approaches will be evaluated by their ability to handle systems at scale. The algorithms and datasets will be disseminated as open-source software. The proposed research makes fundamental contributions to three disciplines: formal methods, machine learning, and control theory; at the same time, it takes fundamental, concrete steps towards broadening participation in computing by making RL-based programming easier and more inclusive.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
The Octatope Abstract Domain for Verification of Neural Networks.
用于验证神经网络的八位位抽象域。
DOI: --
发表时间: 2023
期刊: Formal Methods. FM 2023.
影响因子: --
作者: [Bak, S., Dohmen, T., Subramani, K., Trivedi, A., Velasquez, A., Wojciechowski, P.]
通讯作者: Wojciechowski, P.
DOI: 10.5555/3545946.3599102
发表时间: 2021-10
期刊:
影响因子: --
作者: [Mohammad Afzal;Sankalp Gambhir;Ashutosh Gupta;S. Krishna;Ashutosh Trivedi;Alvaro Velasquez]
通讯作者: Mohammad Afzal;Sankalp Gambhir;Ashutosh Gupta;S. Krishna;Ashutosh Trivedi;Alvaro Velasquez
Mungojerrie: Linear-Time Objectives in Model-Free Reinforcement Learning.
Mungojerrie:无模型强化学习中的线性时间目标。
DOI: --
发表时间: 2023
期刊: Tools and Algorithms for the Construction and Analysis of Systems. TACAS 2023.
影响因子: --
作者: [Hahn, Ernst Moritz, Perez, Mateo, Schewe, Sven, Somenzi, Fabio, Trivedi, Ashutosh, Wojtczak, Dominik]
通讯作者: Wojtczak, Dominik
DOI: --
发表时间: 2022
期刊: Conference on Neural Information Processing Systems (NeurIPS 2022
影响因子: --
作者: [Hahn, Ernst Moritz, Perez, Mateo, Schewe, Sven, Somenzi, Fabio, Trivedi, Ashutosh, Wojtczak, Dominik]
通讯作者: Wojtczak, Dominik
共 8 条
    Collaborative Research: DASS: Assessing Accountability of Tax Preparation Software Systems
    • 批准号:
      2317207
    • 项目类别:
      Standard Grant
    • 资助金额:
      $22.0万
    • 财政年份:
      2023
    • 负责人:
      Ashutosh Trivedi
    • 依托单位:
    SHF: Small: Omega-Regular Objectives for Model-Free Reinforcement Learning
    • 批准号:
      2009022
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2020
    • 负责人:
      Ashutosh Trivedi
    • 依托单位:
    国内基金
    海外基金
    海桑属杂种区强化(Reinforcement)的检验与遗传基础研究
    • 批准号:
      30800060
    • 项目类别:
      青年科学基金项目
    • 资助金额:
      23.0万元
    • 批准年份:
      2008
    • 负责人:
      周仁超
    • 依托单位: