课题基金 / 基金详情

FMitF: Track I: Synthesis and Verification for Programmatic Reinforcement Learning

FMitF: Track I: Synthesis and Verification for Programmatic Reinforcement Learning
FMITF:第一轨:程序化强化学习的综合和验证
批准号:
2124155
负责人:
He Zhu
金额:
$74.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-10-01 至 2025-09-30

项目摘要

项目成果

He Zhu的其他基金

相似基金

相关文献

中文摘要
翻译
近年来,深度强化学习在众多的决策系统中得到了广泛的应用。然而,在自动驾驶等安全关键领域部署RL技术,引发了人们对其可信度的担忧。RL应用程序的行为可能出人意料,因为这样的系统的特点是状态空间大得难以处理,并且在训练期间只探索了一小部分状态。由于难以识别安全的逻辑概念如何与无法解释的策略网络结构相关,进一步加剧了对深层神经RL策略的信任。该项目的新奇之处在于探索特定于领域的程序,作为值得信赖的RL的新学习表示。它允许RL系统像普通编程系统一样进行解释、正式验证和调试。该项目的影响是形成新的方法来设计高度可解释和可验证的RL系统,能够做出可靠的决策。该项目的主要贡献是一个程序性的RL框架,用于合成满足形式正确性规范和量化性能目标的策略。首先,该项目开发了程序合成技术,将原始的RL策略组合成复合的和可解释的程序,以解决复杂的RL环境。为了探索新奇的作文,它应用高效的策略梯度方法在语言语法规则的离散空间的连续松弛中进行搜索。其次,该项目在逐个构造的RL循环中协调程序性策略合成和程序验证,将策略性能优化约束在可证明的安全状态空间中。此外,该项目通过将综合和验证框架提升到一阶关系环境,利用关系抽象将程序性策略转移到具有正式保证的看不见的环境,调查了程序性RL的可转移性。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep reinforcement learning (RL) has recently gained tremendous popularity in numerous decision-making systems. However, deploying RL techniques in safety-critical domains such as autonomous driving has drawn concerns about its trustworthiness. RL applications can behave unexpectedly because such systems are characterized by intractably large state spaces and only a small fraction of states are explored during training. Lack of trust in a deep neural RL policy is further exacerbated by the difficulty of identifying how logical notions of safety relate to the uninterpretable policy network structure. The project's novelty is to explore domain-specific programs as a new learning representation for trustworthy RL. It allows an RL system to be interpreted, formally verified, and debugged as an ordinary programming system. The project's impact is to shape new methodologies to design highly interpretable and verifiable RL systems that are capable of making reliable decisions.The main contribution of the project is a programmatic RL framework for synthesizing policies that meet both formal correctness specifications and quantitative performance objectives. Firstly, the project develops program synthesis techniques to compose primitive RL policies into composite and interpretable programs to solve sophisticated RL environments. To explore novel compositions, it applies efficient policy gradient methods to search in a continuous relaxation of the discrete space of language grammar rules. Secondly, the project reconciles programmatic policy synthesis with program verification in a correct-by-construction RL loop, constraining policy performance optimization in provably safe state space. Moreover, this project investigates the transferability of programmatic RL by lifting the synthesis and verification framework to first-order relational environments, leveraging relational abstractions to transfer programmatic policies to unseen environments with formal guarantees.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1145/3488560.3498410
发表时间: 2021-12
期刊: Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining
影响因子: --
作者: [H. Chen;Yunqi Li;Shaoyun Shi;Shuchang Liu;He Zhu;Yongfeng Zhang]
通讯作者: H. Chen;Yunqi Li;Shaoyun Shi;Shuchang Liu;He Zhu;Yongfeng Zhang
DOI: --
发表时间: 2022
期刊:
影响因子: --
作者: [Wenjie Qiu;He Zhu]
通讯作者: Wenjie Qiu;He Zhu
DOI: 10.18653/v1/2022.findings-emnlp.42
发表时间: 2022
期刊:
影响因子: --
作者: [Wenyue Hua;Yongfeng Zhang]
通讯作者: Wenyue Hua;Yongfeng Zhang
DOI: 10.1145/3511808.3557385
发表时间: 2022-08
期刊: Proceedings of the 31st ACM International Conference on Information & Knowledge Management
影响因子: --
作者: [H. Chen;Yunqi Li;He Zhu;Yongfeng Zhang]
通讯作者: H. Chen;Yunqi Li;He Zhu;Yongfeng Zhang
共 7 条
    SHF: Small: Formal Symbolic Reasoning of Deep Reinforcement Learning Systems
    • 批准号:
      2007799
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.0万
    • 财政年份:
      2020
    • 负责人:
      He Zhu
    • 依托单位:
    海外基金