Lyapunov-Regularized Reinforcement Learning for Power System Transient Stability

Lyapunov-Regularized Reinforcement Learning for Power System Transient Stability
复制标题

DOI:
10.1109/lcsys.2021.3088068
复制
发表时间:
2021-03
影响因子:
3
通讯作者:
Wenqi Cui;Baosen Zhang
Wenqi Cui;Baosen Zhang
中科院分区:
--
文献类型:
--
作者:
Wenqi Cui;Baosen Zhang

文献摘要

相似文献

随着可再生资源的日益整合,电力系统的暂态稳定变得越来越重要。这些资源减少了机械惯性,但也增加了频率响应的灵活性。也就是说,它们的电力电子接口可以执行几乎任意的控制律。为了设计这些控制器,强化学习(RL)已成为寻找由神经网络参数表示的最优非线性控制策略的有效方法。一个关键的挑战是强制要求学习的控制器必须稳定。本文提出了一种求解有损网络暂态稳定最优频率控制的Lyapunov正则化RL方法。由于缺乏解析的Lyapunov函数,我们学习了一个由神经网络参数化的Lyapunov函数。损耗是针对物理电源系统而专门设计的。然后利用学习的神经Lyapunov函数作为正则化,通过惩罚违反Lyapunov条件的行为来训练神经网络控制器。算例分析表明,引入Lyapunov正则化后,控制器能够镇定且损失较小。
Transient stability of power systems is becoming increasingly important because of the growing integration of renewable resources. These resources lead to a reduction in mechanical inertia but also provide increased flexibility in frequency responses. Namely, their power electronic interfaces can implement almost arbitrary control laws. To design these controllers, reinforcement learning (RL) has emerged as a powerful method in searching for optimal non-linear control policy parameterized by neural networks. A key challenge is to enforce that a learned controller must be stabilizing. This letter proposes a Lyapunov regularized RL approach for optimal frequency control for transient stability in lossy networks. Because the lack of an analytical Lyapunov function, we learn a Lyapunov function parameterized by a neural network. The losses are specially designed with respect to the physical power system. The learned neural Lyapunov function is then utilized as a regularization to train the neural network controller by penalizing actions that violate the Lyapunov conditions. Case study shows that introducing the Lyapunov regularization enables the controller to be stabilizing and achieve smaller losses.