CPS: Medium: Collaborative Research: Provably Safe and Robust Multi-Agent Reinforcement Learning with Applications in Urban Air Mobility
CPS: Medium: Collaborative Research: Provably Safe and Robust Multi-Agent Reinforcement Learning with Applications in Urban Air Mobility
批准号:
2312094
负责人:
Quanquan Gu
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2026-05-31
中文摘要
该网络物理系统(CPS)项目旨在为可扩展的多智能体规划和控制设计理论和算法,以支持高吞吐量、不确定和动态环境中安全关键的自主eVTOL飞机。城市空中交通(UAM)是一种新兴的航空运输模式,其中电动垂直起降(eVTOL)飞机将安全有效地在城市区域内运输乘客和货物。白宫、美国国家工程院和美国国会的指导鼓励UAM进行基础研究,以保持美国在该领域的全球领导地位。UAM的成功将取决于安全和强大的多智能体自主性,以扩大高吞吐量城市空中交通的运营。基于学习的技术,如深度强化学习和多智能体强化学习的发展,以支持这些eVTOL车辆的规划和控制。然而,在多智能体自治UAM应用中,为这些基于学习的神经网络在环模型提供理论上的安全性和鲁棒性保证是一个主要的挑战。在这个项目中,研究人员将与政府和行业合作伙伴合作,开展基于用例的基础研究,重点是促进不同背景学生的人工智能、机器学习和自主性的安全性和可靠性。本项目的技术目标包括:(1)单智能体强化学习的安全性和鲁棒性:为了解决“安全临界”UAM挑战,pi规划了单智能体强化学习的最小-最大优化,以形式化地建立足够的安全裕度,约束强化学习将安全作为状态和动作空间的物理约束,以及新颖的谨慎强化学习,使用变分策略梯度以最小的分布风险规划最安全的飞机轨迹;(2)多智能体强化学习的安全性和鲁棒性:为了解决“异构智能体和可扩展性”的挑战,提出了一种新的联邦强化学习框架,其中中央智能体与分散的安全智能体协调以提高交通吞吐量,同时保证安全,以及一种适应不同数量分散飞机的扩展机制;(3)从模拟到现实世界的安全性和鲁棒性:为了解决“高维和环境不确定性”的挑战,研究人员将重点关注智能体在分布转移和从模拟到现实世界的快速适应下的策略鲁棒性。具体来说,以价值为目标的模型学习,以纳入飞机和环境物理等领域知识,并计划在RL模型在线部署用于飞行测试或执行后建立安全的适应机制。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This Cyber-Physical Systems (CPS) project aims at designing theories and algorithms for scalable multi-agent planning and control to support safety-critical autonomous eVTOL aircraft in high-throughput, uncertain and dynamic environments. Urban Air Mobility (UAM) is an emerging air transportation mode in which electrical vertical take-off and landing (eVTOL) aircraft will safely and efficiently transport passengers and cargo within urban areas. Guidance from the White House, the National Academy of Engineering, and the US Congress has encouraged fundamental research in UAM to maintain the US global leadership in this field. The success of UAM will depend on the safe and robust multi-agent autonomy to scale up the operations to high-throughput urban air traffic. Learning-based techniques such as deep reinforcement learning and multi-agent reinforcement learning are developed to support planning and control for these eVTOL vehicles. However, there is a major challenge to provide theoretical safety and robustness guarantees for these learning-based neural network in-the-loop models in multi-agent autonomous UAM applications. In this project, the researchers will collaborate with committed government and industry partners on the use-case-inspired fundamental research, with a focus on promoting safety and reliability of AI, machine learning and autonomy in students with diverse backgrounds. The technical objectives of this project include (1) Safety and Robustness of Single-Agent Reinforcement Learning: in order to address the “safety critical” UAM challenge, the PIs plan the min-max optimization for single agent reinforcement learning to formally build sufficient safety margin, constrained reinforcement learning to formulate safety as physical constraints in state and action spaces, and the novel cautious reinforcement learning that uses variational policy gradient to plan the safest aircraft trajectory with minimum distributional risk; (2) Safety and Robustness of Multi-Agent Reinforcement Learning: in order to address the “heterogeneous agents and scalability” challenge, a novel federated reinforcement learning framework where a central agent coordinates with decentralized safe agents to improve traffic throughput while guaranteeing safety, and a scaling mechanism to accommodate a varying number of decentralized aircraft; (3) Safety and Robustness from Simulations to the Real World: in order to address the “high-dimensionality and environment uncertainty” challenge, the researchers will focus on the agents’ policy robustness under distribution shift and fast adaptation from simulation to the real world. Specifically, value-targeted model learning to incorporate domain knowledge such as the aircraft and environment physics, and a safe adaptation mechanism after the RL model is deployed online for flight testing or execution is planned.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(1)
专著(0)
科研奖励(0)
会议论文
DOI:
10.48550/arxiv.2310.00968
发表时间:
2023-10
期刊:
ArXiv
影响因子:
--
作者:
[Qiwei Di;Tao Jin;Yue Wu;Heyang Zhao;Farzad Farnoud;Quanquan Gu]
通讯作者:
Qiwei Di;Tao Jin;Yue Wu;Heyang Zhao;Farzad Farnoud;Quanquan Gu
Collaborative Research: Towards the Foundation of Approximate Sampling-Based Exploration in Sequential Decision Making
-
批准号:2323113
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2023
-
负责人:Quanquan Gu
-
依托单位:
III: Small: Towards the Foundations of Training Deep Neural Networks: New Theory and Algorithms
-
批准号:2008981
-
项目类别:Continuing Grant
-
资助金额:$50.0万
-
财政年份:2020
-
负责人:Quanquan Gu
-
依托单位:
CIF: Small: Collaborative Research: Rank Aggregation with Heterogeneous Information Sources: Efficient Algorithms and Fundamental Limits
-
批准号:1911168
-
项目类别:Standard Grant
-
资助金额:$25.0万
-
财政年份:2019
-
负责人:Quanquan Gu
-
依托单位:
III: Small: Collaborative Research: High-Dimensional Machine Learning Methods for Personalized Cancer Genomics
-
批准号:1903202
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2018
-
负责人:Quanquan Gu
-
依托单位:
BIGDATA: F: Collaborative Research: Taming Big Networks via Embedding
-
批准号:1855099
-
项目类别:Standard Grant
-
资助金额:$49.99万
-
财政年份:2018
-
负责人:Quanquan Gu
-
依托单位:
CAREER: Scaling Up Knowledge Discovery in High-Dimensional Data Via Nonconvex Statistical Optimization
-
批准号:1906169
-
项目类别:Continuing Grant
-
资助金额:$50.6万
-
财政年份:2018
-
负责人:Quanquan Gu
-
依托单位:
BIGDATA: F: Collaborative Research: Taming Big Networks via Embedding
-
批准号:1741342
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2018
-
负责人:Quanquan Gu
-
依托单位:
III: Small: Collaborative Learning with Incomplete and Noisy Knowledge
-
批准号:1904183
-
项目类别:Standard Grant
-
资助金额:$35.09万
-
财政年份:2018
-
负责人:Quanquan Gu
-
依托单位:
III: Small: Collaborative Research: High-Dimensional Machine Learning Methods for Personalized Cancer Genomics
-
批准号:1717206
-
项目类别:Continuing Grant
-
资助金额:$30.0万
-
财政年份:2017
-
负责人:Quanquan Gu
-
依托单位:
CAREER: Scaling Up Knowledge Discovery in High-Dimensional Data Via Nonconvex Statistical Optimization
-
批准号:1652539
-
项目类别:Continuing Grant
-
资助金额:$51.58万
-
财政年份:2017
-
负责人:Quanquan Gu
-
依托单位:
III: Small: Collaborative Learning with Incomplete and Noisy Knowledge
-
批准号:1618948
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2016
-
负责人:Quanquan Gu
-
依托单位:
海外基金