Impact of Sensor and Actuator Clock Offsets on Reinforcement Learning

Impact of Sensor and Actuator Clock Offsets on Reinforcement Learning
复制标题

DOI:
10.23919/acc53348.2022.9867823
复制
发表时间:
2022-06
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Filippos Fotiadis;Aris Kanellopoulos;K. Vamvoudakis;J. Hugues
Filippos Fotiadis;Aris Kanellopoulos;K. Vamvoudakis;J. Hugues
中科院分区:
其他
文献类型:
--
作者:
Filippos Fotiadis;Aris Kanellopoulos;K. Vamvoudakis;J. Hugues

文献摘要

被引文献

相似文献

在这项工作中,我们研究了传感器执行器时钟偏移对启用强化学习 (RL) 的网络物理系统的影响。特别是,我们考虑一种离策略 RL 算法,该算法从系统的传感器和执行器接收数据,并使用它们来近似所需的最优控制策略。然而,由于时序不匹配,从这些系统组件获得的控制状态数据不一致,因此产生了 RL 鲁棒性的问题。经过广泛的分析,我们表明强化学习确实在 epsilon-delta 意义上保持了鲁棒性;假设传感器-执行器时钟偏移不是任意大,并且行为控制输入满足 Lipschitz 连续性条件,RL 会收敛到接近所需的最优控制策略。在双连杆机械臂上进行了仿真,阐明并验证了理论结果。
In this work, we investigate the effect of sensor-actuator clock offsets on reinforcement learning (RL) enabled cyber-physical systems. In particular, we consider an off-policy RL algorithm that receives data both from the system’s sensors and actuators, and uses them to approximate a desired optimal control policy. Nevertheless, owing to timing mismatches, the control-state data obtained from these system components are inconsistent, hence creating the question of how robust RL will be. After an extensive analysis, we show that RL does retain its robustness, in an epsilon-delta sense; given that the sensor-actuator clock offsets are not arbitrarily large, and that the behavioral control input satisfies a Lipschitz continuity condition, RL converges epsilon-close to the desired optimal control policy. Simulations are carried out on a two-link manipulator, which clarify and verify theoretical findings.