EAGER: Real-Time: Formal Reinforcement Learning Methods for the Design of Safety-critical Autonomous Systems
EAGER: Real-Time: Formal Reinforcement Learning Methods for the Design of Safety-critical Autonomous Systems
批准号:
1839842
负责人:
Rahul Jain
金额:
$28.61万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-04-01 至 2022-03-31
中文摘要
这个探索性研究早期概念拨款(EAGER)项目通过整合形式化方法和数据强化学习,采用全新的第一性原理方法来设计安全关键型自主系统。最近几起涉及半自动驾驶汽车的引人注目的交通事故引发了人们的质疑,即目前以人工智能(AI)为中心的方法是否能让我们达到4级或5级自动驾驶,即实现在所有驾驶场景中性能与人类驾驶员相当的全自动驾驶汽车。另一方面,基于验证和综合的形式化方法的方法可以提供安全保证,但在有效推理不确定性和数据驱动模型的正确性方面存在困难。这个项目将结合这两个看似不相容的范例来设计自主系统。它将使用无模型强化学习算法从半自动驾驶汽车的驾驶数据中学习。它将采用基于模型的方法进行系统设计、验证和综合,以在高度不确定的情况下提供可证明的安全操作。将建立一个自动驾驶测试平台,利用来自缩放车辆模型的人类驾驶数据,利用模仿和逆强化学习算法推断安全控制策略。该研究与具有重大社会意义的智能自主交通系统科学相关。实验测试平台将用于为本科生和K-12外展工作提供实践研究经验。特别是,该项目将开发一个框架,用于以信号时间逻辑表示的安全性和性能规范的最佳控制综合。然后将车辆和行人运动学纳入通过概率计算树逻辑指定的非确定性/概率转移模型中。最后,它将通过学习安全人类驾驶员的痕迹,为受安全规范和复杂时间目标约束的部分观察动态模型开发正式的强化学习方法。该项目的一个关键技术贡献将是开发新的正式强化学习方法,这些方法可能在广泛的应用中有用,其中我们必须通过从数据中学习来合成满足某些安全规范的最优控制器。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This EArly-Concept Grant for Exploratory Research (EAGER) project takes a clean-slate first-principles approach to the design of safety-critical autonomous systems by integrating formal methods and reinforcement learning from data. Several recent high-profile traffic incidents involving semi-autonomous vehicles have raised questions about whether current artificial intelligence (AI)-centered methods can ever lead us to Level 4 or 5 autonomy, i.e., to the realization of fully-autonomous vehicles with performance equivalent to a human driver in all driving scenarios. On the other hand, approaches rooted in formal methods for verification and synthesis can provide safety guarantees but have difficulty in efficiently reasoning about uncertainty and the correctness of data-driven models. This project will combine these two, seemingly incompatible, paradigms for designing autonomous systems. It will use model-free reinforcement learning algorithms to learn from semi-autonomous vehicle driving data. It will adopt model-based methods for system design, verification, and synthesis to offer provably safe operation in highly uncertain scenarios. An AutoDrive testbed will be set up where human driving data from scaled vehicular models will be leveraged to infer safe control policies using imitation and inverse reinforcement learning algorithms. The research is relevant to the science of intelligent autonomous transportation systems with significant societal implications. The experimental testbed will be used to provide hands-on research experience to undergraduate students and for K-12 outreach efforts.In particular, the project will develop a framework for optimal control synthesis for safety and performance specification expressed in signal temporal logic. It will then incorporate vehicular and pedestrian kinematics in non-deterministic/probabilistic transition models specified via probabilistic computation tree logic. Finally, it will develop formal reinforcement learning methods for partially observed dynamic models subject to safety specifications and complex temporal goals by learning from traces of safe human drivers. One key technical contribution of the project will be development of new formal reinforcement learning methods that may be useful in a broad array of applications wherein we must synthesize optimal controllers that satisfy certain safety specifications by learning from data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Model-Free Reinforcement Learning for Optimal Control of Markov Decision Processes Under Signal Temporal Logic Specifications
信号时序逻辑规范下马尔可夫决策过程最优控制的无模型强化学习
DOI:
10.1109/cdc45484.2021.9683444
发表时间:
2021
期刊:
2021 60th IEEE Conference on Decision and Control (CDC
影响因子:
--
作者:
[Kalagarla, Krishna C., Jain, Rahul, Nuzzo, Pierluigi]
通讯作者:
Nuzzo, Pierluigi
Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes”, Proc. ICML 2020. (arXiv:1910.07072)
无限视野平均奖励马尔可夫决策过程中的无模型强化学习,Proc。
DOI:
--
发表时间:
2020
期刊:
Proceedings of the 37th International Conference on Machine Learning
影响因子:
--
作者:
[Chen-Yu Wei, Mehdi Jafarnia-Jahromi]
通讯作者:
Chen-Yu Wei, Mehdi Jafarnia-Jahromi
Online Learning-based Real-time Control of Unknown Autonomous Systems
-
批准号:1810447
-
项目类别:Standard Grant
-
资助金额:$33.0万
-
财政年份:2018
-
负责人:Rahul Jain
-
依托单位:
AF: Small: A New Approach to Analysis and Design of Algorithms for Stochastic Control and Optimization
-
批准号:1817212
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2018
-
负责人:Rahul Jain
-
依托单位:
Collaborative Research: Smarter Markets for a Smarter Grid: Pricing Randomness, Flexibility and Risk
-
批准号:1611574
-
项目类别:Standard Grant
-
资助金额:$22.5万
-
财政年份:2016
-
负责人:Rahul Jain
-
依托单位:
CAREER: Network Economics: Theory and Architectures for Incentive-engineered Networks
-
批准号:0954116
-
项目类别:Continuing Grant
-
资助金额:$42.5万
-
财政年份:2010
-
负责人:Rahul Jain
-
依托单位:
NetSE: Small: Cooperation and Incentives in Communication and Social Networks
-
批准号:0917410
-
项目类别:Continuing Grant
-
资助金额:$44.25万
-
财政年份:2009
-
负责人:Rahul Jain
-
依托单位:
国内基金
海外基金
Immuno-Real Time PCR法精确定量血清MG7抗原及在早期胃癌预警中的价值
-
批准号:30600737
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2006
-
负责人:陈峥
-
依托单位:
无色ReAl3(BO3)4(Re=Y,Lu)系列晶体紫外倍频性能与器件研究
-
批准号:60608018
-
项目类别:青年科学基金项目
-
资助金额:28.0万元
-
批准年份:2006
-
负责人:叶宁
-
依托单位: