FMitF: Track I: Safe Multi-Agent Reinforcement Learning with Shielding
FMitF: Track I: Safe Multi-Agent Reinforcement Learning with Shielding
批准号:
2319500
负责人:
Stavros Tripakis
金额:
$75.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30
中文摘要
该项目结合了形式化方法(FM)和强化学习(RL)的专业知识,开发了一种构建安全多智能体强化学习(MARL)系统的新方法。强化学习方法擅长解决现实环境中的复杂任务(例如,机器人学习在未知环境中导航),但不能提供安全的“硬”保证(例如,机器人保证永远不会相互碰撞)。正式的方法提供了严格的安全保证,但很难扩展到现实世界的设置。该项目旨在将两者的优点结合起来,以设计出能够学习解决现实环境中复杂任务的方法,同时确保安全。该项目的创新之处在于,将FM的屏蔽合成技术和RL的定向勘探技术结合在一起,形成了一种以安全为中心的新方法,并将其应用到一个工具套件中,并对一组基准进行了评估。该项目的影响在于改变RL系统的开发和部署方式,使其能够用于安全关键环境。更广泛的影响包括扩大对不同群体的研究参与,并让本科生参与进来。项目的关键概念是安全防护罩和安全教练。在安全至关重要的情况下(无论是在训练期间还是在执行期间)使用安全防护罩。屏蔽层通过拦截(和修改)代理可能不安全的操作来防止违反安全的行为。当安全违规行为可以被容忍时(例如,在虚拟训练或模拟执行中),将使用安全教练。教练通过鼓励代理犯错误并从中吸取教训来进行安全培训。屏蔽是一个已知的概念,但尚未在最常见的MARL设置(分散执行或部分可观察性)中进行研究。教练是本项目引入的一个新概念,它将在解决任务的同时教导agent注意安全。该项目将开发(1)新的形式化方法和概念,特别是分散的屏蔽综合和安全教练;(2)新的MARL技术,特别是安全导向的探索、安全培训和将安全信息直接硬连接到代理策略中;(3)模型学习和抽象细化到MARL设置的新应用。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This project combines expertise from formal methods (FM) and reinforcement learning (RL) to develop a novel methodology for building safe multi-agent RL (MARL) systems. RL methods are good at solving complex tasks in real-world environments (e.g., robots learning to navigate unknown environments), but cannot provide "hard" guarantees on safety (e.g., robots guaranteed to never collide with each other). Formal methods provide rigorous safety guarantees, but are difficult to scale to real-world settings. This project seeks to combine the best of both worlds in order to devise methods that are capable of learning to solve complex tasks in real-world environments, while at the same time ensuring safety. The project's novelties are a combination of techniques such as shield synthesis from FM and directed exploration from RL into a novel safety-focused methodology, as well as its implementation into a tool suite and its evaluation on a set of benchmarks. The project's impacts are in transforming the way RL systems are developed and deployed so that they can be used in safety-critical settings. Broader impacts include broadening participation in research to diverse groups and involving undergraduate students.Key concepts of the project are safety shields and safety coaches. Safety shields are to be used in safety-critical situations, where safety is paramount (either during training or during execution). Shields prevent safety violations by intercepting (and modifying) potentially unsafe actions by the agents. Safety coaches are to be used when safety violations can be tolerated (e.g., in virtual training or simulated execution). Coaches train for safety by encouraging agents to make mistakes, and to learn from them. Shields are a known concept, but have not been studied in the most common MARL settings—decentralized execution or partial observability. Coaches are a novel concept introduced in this project, which will teach agents to be safe while solving the task. The project will develop (1) new formal methods and concepts, specifically decentralized shield synthesis and safety coaches, (2) new MARL techniques, specifically, safety-directed exploration, training for safety, and hardwiring safety information directly into agent policies, and (3) novel applications of model learning and abstraction refinement to the MARL setting.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SaTC: CORE: Medium: Collaborative: Bridging the Gap between Protocol Design and Implementation through Automated Mapping
-
批准号:1801546
-
项目类别:Continuing Grant
-
资助金额:$73.6万
-
财政年份:2018
-
负责人:Stavros Tripakis
-
依托单位:
CPS: Breakthrough: Compositional System Modeling with Interfaces (COSMOI)
-
批准号:1329759
-
项目类别:Standard Grant
-
资助金额:$49.89万
-
财政年份:2013
-
负责人:Stavros Tripakis
-
依托单位:
海外基金