FMitF: Track I: Synthesis and Verification for Programmatic Reinforcement Learning
FMitF: Track I: Synthesis and Verification for Programmatic Reinforcement Learning
批准号:
2124155
负责人:
He Zhu
金额:
$74.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-10-01 至 2025-09-30
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Deep reinforcement learning (RL) has recently gained tremendous popularity in numerous decision-making systems. However, deploying RL techniques in safety-critical domains such as autonomous driving has drawn concerns about its trustworthiness. RL applications can behave unexpectedly because such systems are characterized by intractably large state spaces and only a small fraction of states are explored during training. Lack of trust in a deep neural RL policy is further exacerbated by the difficulty of identifying how logical notions of safety relate to the uninterpretable policy network structure. The project's novelty is to explore domain-specific programs as a new learning representation for trustworthy RL. It allows an RL system to be interpreted, formally verified, and debugged as an ordinary programming system. The project's impact is to shape new methodologies to design highly interpretable and verifiable RL systems that are capable of making reliable decisions.The main contribution of the project is a programmatic RL framework for synthesizing policies that meet both formal correctness specifications and quantitative performance objectives. Firstly, the project develops program synthesis techniques to compose primitive RL policies into composite and interpretable programs to solve sophisticated RL environments. To explore novel compositions, it applies efficient policy gradient methods to search in a continuous relaxation of the discrete space of language grammar rules. Secondly, the project reconciles programmatic policy synthesis with program verification in a correct-by-construction RL loop, constraining policy performance optimization in provably safe state space. Moreover, this project investigates the transferability of programmatic RL by lifting the synthesis and verification framework to first-order relational environments, leveraging relational abstractions to transfer programmatic policies to unseen environments with formal guarantees.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(7)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1145/3488560.3498410
发表时间:
2021-12
期刊:
Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining
影响因子:
--
作者:
[H. Chen;Yunqi Li;Shaoyun Shi;Shuchang Liu;He Zhu;Yongfeng Zhang]
通讯作者:
H. Chen;Yunqi Li;Shaoyun Shi;Shuchang Liu;He Zhu;Yongfeng Zhang
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Wenjie Qiu;He Zhu]
通讯作者:
Wenjie Qiu;He Zhu
DOI:
10.18653/v1/2022.findings-emnlp.42
发表时间:
2022
期刊:
影响因子:
--
作者:
[Wenyue Hua;Yongfeng Zhang]
通讯作者:
Wenyue Hua;Yongfeng Zhang
DOI:
10.1145/3511808.3557385
发表时间:
2022-08
期刊:
Proceedings of the 31st ACM International Conference on Information & Knowledge Management
影响因子:
--
作者:
[H. Chen;Yunqi Li;He Zhu;Yongfeng Zhang]
通讯作者:
H. Chen;Yunqi Li;He Zhu;Yongfeng Zhang
DOI:
--
发表时间:
2021
期刊:
影响因子:
--
作者:
[Guofeng Cui;He Zhu]
通讯作者:
Guofeng Cui;He Zhu
共 7 条
SHF: Small: Formal Symbolic Reasoning of Deep Reinforcement Learning Systems
-
批准号:2007799
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2020
-
负责人:He Zhu
-
依托单位:
海外基金