Stability Constrained Reinforcement Learning for Decentralized Real-Time Voltage Control

Stability Constrained Reinforcement Learning for Decentralized Real-Time Voltage Control
复制标题

DOI:
10.1109/tcns.2023.3338240
复制
发表时间:
2022-09
影响因子:
4.2
通讯作者:
Jie Feng;Yuanyuan Shi;Guannan Qu;S. Low;Anima Anandkumar;A. Wierman
Jie Feng;Yuanyuan Shi;Guannan Qu;S. Low;Anima Anandkumar;A. Wierman
中科院分区:
计算机科学3区
文献类型:
--
作者:
Jie Feng;Yuanyuan Shi;Guannan Qu;S. Low;Anima Anandkumar;A. Wierman

文献摘要

被引文献

相似文献

深度强化学习被认为是解决电力系统实时控制挑战的一种有前途的工具。然而,由于缺乏明确的稳定性和安全保证,它在现实世界电力系统中的部署受到阻碍。在本文中,我们提出了一个稳定性约束的强化学习(RL)方法的实时电压控制,保证系统的稳定性,在政策学习和部署的学习政策。我们的方法的关键思想是一个显式构造的李雅普诺夫函数,导致稳定政策的充分结构条件,即,单调递减策略保证稳定性。我们将这种结构约束与RL,通过参数化每个局部电压控制器使用单调神经网络,从而确保稳定性约束满足设计。我们证明了我们的方法在单相和三相IEEE测试馈线,所提出的方法可以减少超过25%的暂态控制成本,平均缩短21.5%的电压恢复时间相比,广泛使用的线性政策,同时始终实现电压稳定的有效性。相比之下,标准RL方法通常无法实现电压稳定性。
Deep reinforcement learning has been recognized as a promising tool to address the challenges in real-time control of power systems. However, its deployment in real-world power systems has been hindered by a lack of explicit stability and safety guarantees. In this paper, we propose a stability-constrained reinforcement learning (RL) method for real-time voltage control, that guarantees system stability both during policy learning and deployment of the learned policy. The key idea underlying our approach is an explicitly constructed Lyapunov function that leads to a sufficient structural condition for stabilizing policies, i.e., monotonically decreasing policies guarantee stability. We incorporate this structural constraint with RL, by parameterizing each local voltage controller using a monotone neural network, thus ensuring the stability constraint is satisfied by design. We demonstrate the effectiveness of our approach in both single-phase and three-phase IEEE test feeders, where the proposed method can reduce the transient control cost by more than 25% and shorten the voltage recovery time by 21.5% on average compared to the widely used linear policy, while always achieving voltage stability. In contrast, standard RL methods often fail to achieve voltage stability.