CPS: Small: Distributed Learning for Control of Cyber-Physical Systems
CPS: Small: Distributed Learning for Control of Cyber-Physical Systems
批准号:
1932011
负责人:
Michael Zavlanos
金额:
$40.75万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2023-09-30
中文摘要
在最先进的网络物理系统(CPS)中,通常使用监督学习或无监督学习来分析数据。然而,在许多这类系统中,规则无法事先确定,而且由于数据的动态性质、数据量大,实际上无法贴标签,以及这些数据是逐段而不是事先全部加入系统的事实,这些数据挖掘技术也不能直接适用。另一方面,CPS的控制通常以基于模型的方式进行,其中期望的控制策略从在设计时导出的高保真系统模型计算,并且可能在运行时更新。然而,这种方法是不适合高度动态CPS,潜在地表示系统的系统,其空间和时间配置可能会迅速改变。事实上,由于配置级别如此之多,几乎不可能使用标准的模型驱动技术导出合适的控制策略。因此,它是至关重要的,以方便设计基于数据的控制器,具有强大的性能保证,在某种程度上,允许自然的运行时控制适应。强化学习(RL)提供了这样一个框架。在强化学习中,智能体在反馈回路中与环境交互,通过采取适当的行动序列来学习最优策略,以优化长期回报。因此,与监督和无监督学习相比,RL在分析流数据,特别是控制系统方面更有效。该项目的目标是开发一个分布式的非策略RL框架,用于控制CPS。分布式RL方法避免了在中央处理器上收集所有信息的脆弱性,通信开销和隐私问题。此外,离政策学习方法显着提高采样效率,并确保更安全的操作。在该项目下开发的分布式RL框架将对CPS的控制产生深远的影响,涉及交通、制造、医疗保健、智慧城市、城市规划等领域,依靠多个传感器进行数据收集和控制。该项目还涉及一个教育议程,重点是K-12,本科和研究生教育。该项目的外联部分侧重于提高大学预科学生对研究和工程职业的潜力和吸引力的认识,该项目的技术目标分为四个方面。第一个推力开发分布式的非策略RL方法,使用线性函数近似的动作值函数。分布式RL算法使用线性函数近似已经提出了政策评估。这一推动力开发了新的RL算法,这些算法也可以改进策略,直到找到控制所需的最优策略。由于为RL问题定义适当的特征向量通常很困难,并且由于线性映射可能无法捕获这些特征之间可能的非线性相互作用,因此第二个推力使用非线性函数近似开发了分布式非策略RL方法,特别是神经网络。第三个推力发展分布式的政策外的演员-评论家的方法。当动作空间很大或连续时,Actor-Critic方法更有效,因为它们使用线性或非线性函数近似来参数化目标策略函数,并学习最佳参数,以便生成的策略映射到每个状态的最佳动作。最后,第四个推力为现代CPS中常见的异步,异构和非平稳数据开发分布式RL方法,其中传感器不观察相同分布的数据,也不同时采样数据。此外,从其采样数据的分布可以随时间而改变。该项目的重点是算法的开发和支持的理论成果。开发的算法在CPS的资源分配问题,特别是在分布式共享车辆调度systems.This奖项的控制模拟评估反映了NSF的法定使命,并已被认为是值得通过使用基金会的智力价值和更广泛的影响审查标准的评估支持。
英文摘要
In state-of-the-art Cyber-Physical-Systems (CPS) supervised learning or unsupervised learning are typically used to analyze data. Nevertheless, in many such systems rules cannot be determined in advance and these data mining techniques are not directly applicable due to the dynamic nature of the data, their large volume that prohibits labelling in practice, and the fact that these data are added to the system piece by piece and not altogether in advance. On the other hand, control of CPS is usually done in a model-based manner, where a desired control policy is computed from a high-fidelity system model that has been derived at design-time, and potentially may be updated at runtime. However, this approach is not suitable for highly dynamical CPS, that potentially represent systems of systems whose spatial and temporal configurations may rapidly change. In fact, with such high number of configuration levels, it is almost impossible to derive suitable control policies using standard model-driven techniques. Consequently, it is critical to facilitate design of data-based controllers, with strong performance guarantees, in a way that allows for natural runtime control adaptation. Reinforcement Learning (RL) provides such a framework. In RL agents interact with the environment in a feedback loop to learn an optimal policy by taking appropriate sequences of actions in order to optimize longterm payoff. As such, RL can be much more efficient compared to supervised and unsupervised learning, in analyzing streaming data and especially in controlling a system. The goal of this project is to develop a distributed off-policy RL framework for the control of CPS. Distributed RL methods avoid the fragility, communication overhead, and privacy concerns of collecting all information at a central processing unit. Moreover, off-policy learning methods significantly improve sampling efficiency and ensure safer operation. The distributed RL framework developed under this project will have a profound impact on the control of CPS, in areas as diverse as transportation, manufacturing, health-care, smart city, urban planning, etc., that rely on multiple sensors for data collection and control. This project also involves an educational agenda focusing on K-12, undergraduate, and graduate level education. The outreach component of this project focuses on improving the pre-college students' awareness of the potential and attractiveness of a research and engineering career.The technical aims of this project are divided into four thrusts. The first thrust develops distributed off-policy RL methods using linear function approximation of the action-value function. Distributed RL algorithms using linear function approximation have been proposed for policy evaluation only. This thrust develops new RL algorithms that can also improve the policy until an optimal policy is found, which is necessary for control. Since defining appropriate feature vectors for RL problems is generally difficult and since linear mappings might not able to capture possibly nonlinear interactions between these features, the second thrust develops distributed off-policy RL methods using nonlinear function approximation, specifically, Neural Networks. The third thrust develops distributed off-policy Actor-Critic methods. When the action space is large or continuous, Actor-Critic methods are much more effective since they parameterize the target policy function using either linear or nonlinear function approximation and learn the optimal parameter so that the resulting policy maps to the optimal action for every state. Finally, the fourth thrust develops distributed RL methods for asynchronous, heterogeneous, and non-stationary data that are common in modern CPS, where sensors do not observe identically distributed data nor do they sample data at the same time. Moreover, the distributions from which data are sampled can change with time. This project focuses on the development of algorithms and supporting theoretical results. The developed algorithms are evaluated in simulation on resource allocation problems in CPS, specifically, on the control of distributed shared vehicle dispatch systems.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(18)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.automatica.2020.109218
发表时间:
2020
期刊:
Automatica
影响因子:
6.4
作者:
[Zhang, Yan, Zavlanos, Michael M.]
通讯作者:
Zavlanos, Michael M.
DOI:
10.48550/arxiv.2203.08957
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
作者:
[Zifan Wang;Yi Shen;M. Zavlanos]
通讯作者:
Zifan Wang;Yi Shen;M. Zavlanos
DOI:
10.1109/tro.2020.2980176
发表时间:
2018-12
期刊:
IEEE Transactions on Robotics
影响因子:
7.8
作者:
[Reza Khodayi-mehr;M. Zavlanos]
通讯作者:
Reza Khodayi-mehr;M. Zavlanos
Transfer Reinforcement Learning under Unobserved Contextual Information
未观察到的上下文信息下的迁移强化学习
DOI:
10.1109/iccps48487.2020.00015
发表时间:
2020
期刊:
2020 ACM/IEEE 11th International Conference on Cyber-Physical Systems (ICCPS
影响因子:
--
作者:
[Zhang, Yan, Zavlanos, Michael M.]
通讯作者:
Zavlanos, Michael M.
DOI:
--
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
作者:
[Panagiotis Vlantis;M. Zavlanos]
通讯作者:
Panagiotis Vlantis;M. Zavlanos
共 16 条
CPS: Medium: Collaborative Research: Human-on-the-Loop Control for Smart Ultrasound Imaging
-
批准号:1837499
-
项目类别:Standard Grant
-
资助金额:$60.0万
-
财政年份:2018
-
负责人:Michael Zavlanos
-
依托单位:
NeTS: Medium: Collaborative Research: Optimal Communication for Faster Sensor Network Coordination
-
批准号:1302284
-
项目类别:Standard Grant
-
资助金额:$26.0万
-
财政年份:2013
-
负责人:Michael Zavlanos
-
依托单位:
NeTS: Synergy: Collaborative Research: Controlling Teams of Autonomous Mobile Beamformers
-
批准号:1239339
-
项目类别:Standard Grant
-
资助金额:$29.4万
-
财政年份:2013
-
负责人:Michael Zavlanos
-
依托单位:
RI: Medium: Collaborative Research: Mobile Microrobot Platform for Advanced Manufacturing Applications
-
批准号:1302283
-
项目类别:Continuing Grant
-
资助金额:$18.45万
-
财政年份:2013
-
负责人:Michael Zavlanos
-
依托单位:
CAREER: Control of Mobile Robot Networks: Integrating the Communication and Physical Domains
-
批准号:1261828
-
项目类别:Continuing Grant
-
资助金额:$34.58万
-
财政年份:2012
-
负责人:Michael Zavlanos
-
依托单位:
CAREER: Control of Mobile Robot Networks: Integrating the Communication and Physical Domains
-
批准号:1054604
-
项目类别:Continuing Grant
-
资助金额:$44.96万
-
财政年份:2011
-
负责人:Michael Zavlanos
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: