Data-efficient Safe Control with Recovery-to-Optimality Guarantees
Data-efficient Safe Control with Recovery-to-Optimality Guarantees
批准号:
2227311
负责人:
Bahare Kiumarsi
金额:
$40.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31
中文摘要
随着近年来学习型系统的快速发展,提高了系统的自主能力,其安全认证变得尤为重要。虽然安全强化学习(RL)算法在自主控制设计方面的最新进展是有希望的,但这些算法只有在稳定的环境和全面而高质量的数据集的情况下才是可靠的。然而,许多系统必须在不可预测的环境中运行,在这种环境下,安全和性能之间可能会出现危险的差异。在这些环境中,需要调整安全和性能规范以适应环境。此外,RL主体必须在真实的数据量和质量下进行学习。当前的RL实践假设可以获得丰富且高质量的数据,并且可以完全观察整个系统的状态。在许多实际系统中,这些假设可能会被违反。该奖项支持为部分可观测系统创建低复杂性安全学习算法的研究,这些系统配备了高效的冲突管理机制,以安全地提供尽可能多的性能。这一进展将对自主系统、机器人、制造、智能电网等领域的应用产生广泛的影响。本研究项目旨在为配备高效冲突管理机制的部分可观测系统开发低复杂度、安全学习的算法。本项目的目标有两个:1)针对部分可观测的动态不确定非线性系统,提出直接数据驱动的后备安全控制策略学习方法。利用L额外样本动力学、概率可逆性和凸提升等概念,可以仅用测量的有噪输入输出数据来学习具有非凸安全集的非线性系统的安全控制策略。2)引入新的合并方法,通过合并学习的备份安全控制策略和支持学习的控制策略来主动管理冲突。这些方法不是在冲突出现时对它们提供被动的快速解决办法,而是使主动的冲突管理能够避免未来的破坏性冲突。对于冲突管理,RL智能体的水平集将根据情况进行调整,使智能体与安全约束保持一致。也就是说,安全形状的价值函数将学习通过考虑相关领域的安全和最优化问题来有效地解决冲突。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
As rapid developments on learning-enabled systems in recent years have been advancing autonomy capabilities of systems, their safety certification becomes exceedingly important. While the recent progress on safe reinforcement learning (RL) algorithms for autonomous control design has been promising, these algorithms are accountable only in stable environments and under the availability of comprehensive and high-quality data sets. However, many systems must operate in unpredictable environments under which dangerous divergence might arise between safety and performance. In these environments, adaptation of safety and performance specifications to the context is required. Besides, RL agent must perform learning under realistic data quantity and quality. Current RL practice assumes availability of rich and high-quality data with full observability of the entire system’s states. These assumptions can be violated in many practical systems. This award supports research to create low-complexity safe learning-enabled algorithms for partially observable systems that are equipped with highly-efficient conflict management mechanisms to deliver as much performance as possible safely. Advances will have broad implications in applications of autonomous systems, robots, manufacturing, smart grids, and more.This research project aims to develop low-complexity, safe learning-enabled algorithms for partially observable systems equipped with highly efficient conflict management mechanisms. The objectives of this project are two-fold: 1) Proposing direct data-driven learning approaches for backup safe control policies in partially observable nonlinear systems with uncertain dynamics. The utilization of concepts such as L-extra sample dynamics, probabilistic contractivity, and convex lifting will enable the learning of safe control policies for nonlinear systems with nonconvex safe sets using only measured noisy input-output data. 2) Introducing novel merging approaches to proactively manage conflicts by merging learned backup safe control policies with learning-enabled control policies. Instead of providing reactive quick fixes to conflicts as they arise, these approaches will enable proactive conflict management to avoid destructive future conflicts. Towards conflict management, the level sets of the RL agent will be adapted to the situation to make the agent align with the safety constraint. That is, safety-shaped value functions will be learned to effectively resolve conflicts by considering safety and optimality concerns across the relevant domains.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
-
批准号:60973026
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2009
-
负责人:鲁道夫
-
依托单位: