课题基金 / 基金详情

CAREER: Temporal Causal Reinforcement Learning and Control for Autonomous and Swarm Cyber-Physical Systems

CAREER: Temporal Causal Reinforcement Learning and Control for Autonomous and Swarm Cyber-Physical Systems
职业:自治和群体网络物理系统的时间因果强化学习和控制
批准号:
2339774
负责人:
Zhe Xu
金额:
$54.98万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2024
资助国家:
美国
项目状态:
未结题
起止时间:
2024-03-01 至 2029-02-28

项目摘要

项目成果

Zhe Xu的其他基金

相似基金

相关文献

中文摘要
翻译
了解行为的根本原因对于明智的决策和防止无效或有偏见的政策是必不可少的。目前,嵌入网络物理系统(CP)中的大多数基于人工智能的学习和控制模块依赖于统计相关性而不是因果关系进行决策。这不仅会导致错误的决策,还会阻碍学习的可解释性,限制可转移性和可伸缩性。这项职业建议旨在弥合因果推理和CPS中不断增长的强化学习(RL)能力之间的差距。所提出的方法对CPS的广泛应用具有变革性,使得自动驾驶汽车、无人机、工业机器人和群体机器人等自主和群体CPS能够更高效地进行决策。NSF职业生涯提案通过利用单智能体、多智能体和群体系统环境下的时序逻辑和因果图的推理能力,提出了一套针对CPS的时序因果RL和控制方法。我们开发的工具将在多个CPS试验床上实施,并与拟议的教育计划相结合。所提出的算法具有以下独特和创新的特点。首先,我们将开发计算高效的工具,在执行RL时从CPS的观察数据和干预数据中发现时间因果知识,以提高采样效率和可转移性。其次,我们将开发多智能体RL方法用于合作、非合作和不完全信息随机博弈环境中的CPS,其中时间因果知识以分布式方式发现,以加快RL。最后,我们将利用智能体级特征和群体级特征(如密度和广义矩)的时序因果推理,开发可扩展的基于RL的群体系统控制方法。该教育计划将通过人工智能辅助的自适应和互动教学、基于时序逻辑的教育游戏、针对时序因果RL的在线互动教育网站设计以及与行业合作伙伴的研讨会和网络研讨会来影响下一代CP和AI工程师和研究人员。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding the root cause of behavior is imperative for informed decision-making and preventing ineffective or biased policies. Currently, most AI-based learning and control modules embedded in cyber-physical systems (CPS) rely on statistical correlation rather than causality for decision-making. This not only results in incorrect decisions but also hinders the interpretability of learning, limiting transferability and scalability. This CAREER proposal aims to bridge the gap between causal inference and the growing capabilities of reinforcement learning (RL) in CPS. The proposed methods are transformative to a wide range of CPS applications, enabling more efficient and effective decision-making processes in autonomous and swarm CPS such as self-driving cars, drones, industrial robots, and swarm robots.This NSF CAREER proposal proposes a set of temporal causal RL and control approaches for CPS by leveraging the reasoning capabilities of temporal logics and causal diagrams in single-agent, multi-agent, and swarm system settings. The tools we develop will be implemented on multiple CPS testbeds and integrated with the proposed education plan. The proposed algorithms have the following unique and innovative features. Firstly, we will develop computationally efficient tools that can discover temporal causal knowledge from both observational and interventional data of a CPS in performing RL to improve the sampling efficiency and transferability. Secondly, we will develop multi-agent RL approaches for CPS in cooperative, non-cooperative, and incomplete information stochastic game environments where temporal causal knowledge is discovered in a distributed way for expediting RL. Lastly, we will develop scalable RL-based control methods for swarm systems utilizing temporal causal reasoning over agent-level features and swarm-level features such as densities and generalized moments. The education plan will impact the next generation of CPS and AI engineers and researchers through AI-assisted adaptive and interactive teaching, temporal-logic-based educational games, online interactive educational website design for temporal causal RL, and workshops and webinars with industrial partners.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CPS: Small: Neuro-Symbolic Learning and Control with High-Level Knowledge Inference
  • 批准号:
    2304863
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2023
  • 负责人:
    Zhe Xu
  • 依托单位:
海外基金