Computationally Efficient Safe Reinforcement Learning for Power Systems

Computationally Efficient Safe Reinforcement Learning for Power Systems
复制标题

DOI:
10.23919/acc53348.2022.9867652
复制
发表时间:
2021-10
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Daniel Tabas;Baosen Zhang
Daniel Tabas;Baosen Zhang
中科院分区:
其他
文献类型:
--
作者:
Daniel Tabas;Baosen Zhang

文献摘要

相似文献

我们提出了一种计算效率高的方法来安全强化学习(RL)的电力系统中的频率调节具有高水平的可变可再生能源资源。该方法利用集合论控制技术来制定基于神经网络的控制策略,该控制策略保证满足安全临界状态约束,而不需要在真实的时间内解决模型预测控制或投影问题。通过利用鲁棒控制不变多面体的特性,我们构建了一个新的,封闭形式的“安全过滤器”,使端到端的安全学习使用任何基于策略梯度的RL算法。然后,我们将安全滤波器与深度确定性策略梯度(DDPG)算法相结合,以调节修改后的9节点电力系统中的频率,并表明学习的策略比鲁棒线性反馈控制技术更具成本效益,同时保持相同的安全保证。我们还表明,所提出的范例优于DDPG增强约束违反处罚。
We propose a computationally efficient approach to safe reinforcement learning (RL) for frequency regulation in power systems with high levels of variable renewable energy resources. The approach draws on set-theoretic control techniques to craft a neural network-based control policy that is guaranteed to satisfy safety-critical state constraints, without needing to solve a model predictive control or projection problem in real time. By exploiting the properties of robust controlled-invariant polytopes, we construct a novel, closed-form "safety-filter" that enables end-to-end safe learning using any policy gradient-based RL algorithm. We then apply the safety filter in conjunction with the deep deterministic policy gradient (DDPG) algorithm to regulate frequency in a modified 9-bus power system, and show that the learned policy is more cost-effective than robust linear feedback control techniques while maintaining the same safety guarantee. We also show that the proposed paradigm outperforms DDPG augmented with constraint violation penalties.