课题基金 / 基金详情

FMitF: Track I: Safe Multi-Agent Reinforcement Learning with Shielding

FMitF: Track I: Safe Multi-Agent Reinforcement Learning with Shielding
FMITF:第一轨:带屏蔽的安全多智能体强化学习
批准号:
2319500
负责人:
Stavros Tripakis
金额:
$75.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30

项目摘要

项目成果

Stavros Tripakis的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
This project combines expertise from formal methods (FM) and reinforcement learning (RL) to develop a novel methodology for building safe multi-agent RL (MARL) systems. RL methods are good at solving complex tasks in real-world environments (e.g., robots learning to navigate unknown environments), but cannot provide "hard" guarantees on safety (e.g., robots guaranteed to never collide with each other). Formal methods provide rigorous safety guarantees, but are difficult to scale to real-world settings. This project seeks to combine the best of both worlds in order to devise methods that are capable of learning to solve complex tasks in real-world environments, while at the same time ensuring safety. The project's novelties are a combination of techniques such as shield synthesis from FM and directed exploration from RL into a novel safety-focused methodology, as well as its implementation into a tool suite and its evaluation on a set of benchmarks. The project's impacts are in transforming the way RL systems are developed and deployed so that they can be used in safety-critical settings. Broader impacts include broadening participation in research to diverse groups and involving undergraduate students.Key concepts of the project are safety shields and safety coaches. Safety shields are to be used in safety-critical situations, where safety is paramount (either during training or during execution). Shields prevent safety violations by intercepting (and modifying) potentially unsafe actions by the agents. Safety coaches are to be used when safety violations can be tolerated (e.g., in virtual training or simulated execution). Coaches train for safety by encouraging agents to make mistakes, and to learn from them. Shields are a known concept, but have not been studied in the most common MARL settings—decentralized execution or partial observability. Coaches are a novel concept introduced in this project, which will teach agents to be safe while solving the task. The project will develop (1) new formal methods and concepts, specifically decentralized shield synthesis and safety coaches, (2) new MARL techniques, specifically, safety-directed exploration, training for safety, and hardwiring safety information directly into agent policies, and (3) novel applications of model learning and abstraction refinement to the MARL setting.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SaTC: CORE: Medium: Collaborative: Bridging the Gap between Protocol Design and Implementation through Automated Mapping
  • 批准号:
    1801546
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $73.6万
  • 财政年份:
    2018
  • 负责人:
    Stavros Tripakis
  • 依托单位:
CPS: Breakthrough: Compositional System Modeling with Interfaces (COSMOI)
  • 批准号:
    1329759
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.89万
  • 财政年份:
    2013
  • 负责人:
    Stavros Tripakis
  • 依托单位:
海外基金