课题基金 / 基金详情

Data-efficient Safe Control with Recovery-to-Optimality Guarantees

Data-efficient Safe Control with Recovery-to-Optimality Guarantees
数据高效的安全控制,并保证恢复最佳性
批准号:
2227311
负责人:
Bahare Kiumarsi
金额:
$40.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-08-15 至 2026-07-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
随着近年来学习系统的快速发展,系统的自主能力不断提高,其安全认证变得非常重要。虽然用于自主控制设计的安全强化学习(RL)算法的最新进展很有希望,但这些算法仅在稳定的环境中以及在全面和高质量数据集的可用性下才能负责。然而,许多系统必须在不可预测的环境中运行,在这种环境下,安全性和性能之间可能会出现危险的分歧。在这些环境中,需要根据具体情况调整安全和性能规范。此外,RL代理必须在现实的数据数量和质量下进行学习。目前的RL实践假设了丰富和高质量的数据的可用性,以及整个系统状态的完全可观测性。这些假设在许多实际系统中可能会被违反。该奖项支持为部分可观察系统创建低复杂度安全学习算法的研究,这些算法配备了高效的冲突管理机制,以安全地提供尽可能多的性能。该研究项目旨在为部分可观测系统开发低复杂度、安全的学习算法,并配备高效的冲突管理机制。该项目的目标是双重的:1)提出直接的数据驱动学习方法的备份安全控制策略的部分可观测的非线性系统的不确定动态。利用的概念,如L-额外的样本动态,概率收缩,凸提升将使学习的安全控制策略的非凸安全集的非线性系统,只使用测量的噪声输入输出数据。2)引入新颖的合并方法,通过将学习的备份安全控制策略与支持学习的控制策略合并来主动管理冲突。这些办法不是在冲突出现时提供反应性的快速解决办法,而是能够进行积极主动的冲突管理,以避免今后发生破坏性的冲突。在冲突管理方面,强化学习智能体的水平集将根据情况进行调整,使智能体与安全约束保持一致。也就是说,安全形价值函数将学习,以有效地解决冲突,通过考虑安全和最优性的关注,在相关领域。这个奖项反映了NSF的法定使命,并已被认为是值得的支持,通过评估使用基金会的知识价值和更广泛的影响审查标准。
英文摘要
As rapid developments on learning-enabled systems in recent years have been advancing autonomy capabilities of systems, their safety certification becomes exceedingly important. While the recent progress on safe reinforcement learning (RL) algorithms for autonomous control design has been promising, these algorithms are accountable only in stable environments and under the availability of comprehensive and high-quality data sets. However, many systems must operate in unpredictable environments under which dangerous divergence might arise between safety and performance. In these environments, adaptation of safety and performance specifications to the context is required. Besides, RL agent must perform learning under realistic data quantity and quality. Current RL practice assumes availability of rich and high-quality data with full observability of the entire system’s states. These assumptions can be violated in many practical systems. This award supports research to create low-complexity safe learning-enabled algorithms for partially observable systems that are equipped with highly-efficient conflict management mechanisms to deliver as much performance as possible safely. Advances will have broad implications in applications of autonomous systems, robots, manufacturing, smart grids, and more.This research project aims to develop low-complexity, safe learning-enabled algorithms for partially observable systems equipped with highly efficient conflict management mechanisms. The objectives of this project are two-fold: 1) Proposing direct data-driven learning approaches for backup safe control policies in partially observable nonlinear systems with uncertain dynamics. The utilization of concepts such as L-extra sample dynamics, probabilistic contractivity, and convex lifting will enable the learning of safe control policies for nonlinear systems with nonconvex safe sets using only measured noisy input-output data. 2) Introducing novel merging approaches to proactively manage conflicts by merging learned backup safe control policies with learning-enabled control policies. Instead of providing reactive quick fixes to conflicts as they arise, these approaches will enable proactive conflict management to avoid destructive future conflicts. Towards conflict management, the level sets of the RL agent will be adapted to the situation to make the agent align with the safety constraint. That is, safety-shaped value functions will be learned to effectively resolve conflicts by considering safety and optimality concerns across the relevant domains.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
固定参数可解算法在平面图问题的应用以及和整数线性规划的关系
  • 批准号:
    60973026
  • 项目类别:
    面上项目
  • 资助金额:
    32.0万元
  • 批准年份:
    2009
  • 负责人:
    鲁道夫
  • 依托单位: