Physics-Model-Regulated Deep Reinforcement Learning Towards Safety & Stability Guarantees

Physics-Model-Regulated Deep Reinforcement Learning Towards Safety & Stability Guarantees
复制标题

DOI:
10.1109/cdc49753.2023.10383560
复制
发表时间:
2023-12
期刊:
2023 62nd IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Hongpeng Cao;Yanbing Mao;Lui Sha;Marco Caccamo
Hongpeng Cao;Yanbing Mao;Lui Sha;Marco Caccamo
中科院分区:
其他
文献类型:
--
作者:
Hongpeng Cao;Yanbing Mao;Lui Sha;Marco Caccamo

文献摘要

相似文献

深度强化学习(DRL)通过从数据中合成控制策略,在解决复杂控制任务方面取得了令人印象深刻的成功。然而,DRL应用于安全关键系统的安全性和稳定性仍然是一个主要关注和具有挑战性的问题。为了解决这个问题,我们提出了Phy-DRL:一种新的物理模型调节的深度强化学习框架。Phy-DRL在两个架构设计上是新颖的:物理模型调节的奖励和残差控制,它集成了基于物理模型的控制和数据驱动的控制。并行设计使Phy-DRL具有数学上可证明的安全性和稳定性保证。最后,通过倒立摆系统验证了Phy-DRL的有效性。实验结果表明,Phy-DRL具有显著的加速训练和扩大奖励的特点。
Deep reinforcement learning (DRL) has demonstrated impressive success in solving complex control tasks by synthesizing control policies from data. However, the safety and stability of applying DRL to safety-critical systems remain a primary concern and challenging problem. To address the problem, we propose the Phy-DRL: a novel physics-model-regulated deep reinforcement learning framework. The Phy-DRL is novel in two architectural designs: a physics-model-regulated reward and residual control, which integrates physics-model-based control and data-driven control. The concurrent designs enable the Phy-DRL the mathematically provable safety and stability guarantees. Finally, the effectiveness of the Phy-DRL is validated by an inverted pendulum system. Additionally, the experimental results demonstrate that the Phy-DRL features remarkably accelerated training and enlarged reward.