课题基金 / 基金详情

CPS: Medium: Collaborative Research: Certifiable reinforcement learning for cyber-physical systems

CPS: Medium: Collaborative Research: Certifiable reinforcement learning for cyber-physical systems
CPS:媒介:协作研究:网络物理系统的可认证强化学习
批准号:
1836932
负责人:
Samuel Coogan
金额:
$32.97万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2018
资助国家:
美国
项目状态:
已结题
起止时间:
2018-09-15 至 2023-08-31

项目摘要

项目成果

Samuel Coogan的其他基金

相似基金

相关文献

中文摘要
翻译
我们建议推广和验证强化学习算法的性能,用于控制网络物理系统(CPS)。一般来说,应用于物理系统的强化学习涉及从数据中进行预测,以控制系统使性能标准达到极值。该项目将特别侧重于开发适用于混合和多代理控制系统的理论和算法,即具有连续和离散元素的系统以及具有多个决策代理的系统,这些系统在CPS中无处不在,跨越时空尺度和应用领域。强化学习算法还不够成熟,不能保证应用于CPS控制时的性能。鉴于这些局限性,本项目旨在为强化学习算法的认证奠定理论和计算基础,使其能够以高置信度在社会中部署。本项目将认证在具有非经典动力学和非经典成本的系统中计算最优控制策略的强化学习算法。为了实现这一目标,我们将推广收敛算法最初设计的纯连续系统应用于混合控制系统的状态进行混合的离散和连续的过渡。此外,我们的具体目标是确保这种方法适用于社会规模的CPS,其中多个代理,其中一些可能是人类,直接与CPS互动。这些算法将在三个测试平台上进行实验验证,这些测试平台代表了CPS中出现的一系列混合动力和多智能体现象。第一个测试平台将通过仿真测试我们的算法在社会规模的交通流网络上的性能。第二个测试平台将考虑由空中和地面移动的机器人组成的异构团队,与人类合作伙伴合作,对桥梁和隧道等基础设施的大规模复制品进行施工,检查和维护任务。第三个测试平台将研究人类个体与执行动态运动和操纵行为的远程遥控机器人之间的闭环交互。该项目还将与技术政策专家共同组织一次跨学科研讨会,其结果将成为由PI举办的跨学科多校区研究生级别研讨会的基础。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估而被认为值得支持。
英文摘要
We propose to generalize and certify the performance of reinforcement learning algorithms for control of cyber-physical systems (CPS). Broadly speaking, reinforcement learning applied to physical systems is concerned with making predictions from data to control the system to extremize a performance criterion. The project will particularly focus on developing theory and algorithms applicable to hybrid and multi-agent control systems, that is, systems with continuous and discrete elements and systems with multiple decision-making agents, which are ubiquitous in CPS across spatiotemporal scales and application domains. Reinforcement learning algorithms are not yet mature enough to guarantee performance when applied to control of CPS. In light of these limitations, this project aims to lay the theoretical and computational foundation to certify reinforcement learning algorithms so that they may be deployed in society with high confidence.This project will certify reinforcement learning algorithms that compute optimal control policies in systems with non-classical dynamics and non-classical costs. To achieve this goal, we will generalize convergent algorithms originally designed for purely continuous systems to apply in hybrid control systems whose states undergo a mixture of discrete and continuous transitions. Moreover, we specifically aim to ensure this approach is applicable to societal-scale CPS in which multiple agents, some of which may be humans, interact directly with the CPS. These algorithms will be experimentally validated on three testbeds that represent a range of hybrid and multi-agent phenomena that arise in CPS. The first testbed will test the performance of our algorithms on societal-scale traffic flow networks via simulation. The second testbed will consider heterogeneous teams of aerial and terrestrial mobile robots collaborating with human partners to perform construction, inspection, and maintenance tasks on scale facsimiles of infrastructure like bridges, and tunnels. The third testbed will study the closed-loop interaction between individual humans and remote, teleoperated robots that perform dynamic locomotion and manipulation behaviors. This project will also co-organize an interdisciplinary workshop with technology policy experts, the results of which will form the basis for an interdisciplinary multi-campus graduate-level seminar run by the PIs.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(14)
专著(0)
科研奖励(0)
会议论文
Estimating High Probability Reachable Sets using Gaussian Processes
使用高斯过程估计高概率可达集
DOI: 10.1109/cdc45484.2021.9682962
发表时间: 2021
期刊: IEEE Conference on Decision and Control
影响因子: --
作者: [Cao, Michael E., Bloch, Matthieu, Coogan, Samuel]
通讯作者: Coogan, Samuel
DOI: 10.48550/arxiv.2208.03889
发表时间: 2022-08
期刊: ArXiv
影响因子: --
作者: [Saber Jafarpour;A. Davydov;Matthew Abate;F. Bullo;S. Coogan]
通讯作者: Saber Jafarpour;A. Davydov;Matthew Abate;F. Bullo;S. Coogan
DOI: --
发表时间: 2022
期刊: IEEE Control Systems Letters
影响因子: 3
作者: [Abate, Matthew, Coogan, Samuel]
通讯作者: Coogan, Samuel
DOI: 10.1109/iros45743.2020.9341190
发表时间: 2020-03
期刊: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子: --
作者: [Mohit Srinivasan;A. Dabholkar;S. Coogan;P. Vela]
通讯作者: Mohit Srinivasan;A. Dabholkar;S. Coogan;P. Vela
12
    CAREER: Correct-By-Design Control of Traffic Flow Networks
    • 批准号:
      1749357
    • 项目类别:
      Standard Grant
    • 资助金额:
      $50.01万
    • 财政年份:
      2018
    • 负责人:
      Samuel Coogan
    • 依托单位:
    海外基金