Stable Reinforcement Learning for Optimal Frequency Control: A Distributed Averaging-Based Integral Approach

Stable Reinforcement Learning for Optimal Frequency Control: A Distributed Averaging-Based Integral Approach
复制标题

DOI:
10.1109/ojcsys.2022.3202202
复制
发表时间:
2022-05
期刊:
IEEE Open Journal of Control Systems
影响因子:
--
通讯作者:
Yan Jiang;Wenqi Cui;Baosen Zhang;Jorge Cort'es
Yan Jiang;Wenqi Cui;Baosen Zhang;Jorge Cort'es
中科院分区:
其他
文献类型:
--
作者:
Yan Jiang;Wenqi Cui;Baosen Zhang;Jorge Cort'es

文献摘要

相似文献

频率控制对电力系统的可靠运行起着关键作用。它通常以分级方式执行,首先快速稳定频率偏差,然后缓慢恢复标称频率。然而,随着发电组合从同步发电机转移到可再生能源,电力系统由于惯性损失而经历更大和更快的频率波动,这对频率稳定性产生不利影响。这激发了在快速时间尺度内联合解决频率退化和经济效率的算法的积极研究,其中基于分布式平均的积分(DAI)控制是一个值得注意的算法,其将可控功率注入设置成与频率偏差和经济效率低下信号的积分成正比。然而,DAI通常不考虑功率扰动后系统的瞬态性能,并且仅限于二次操作成本函数。本文的目的是利用非线性最优控制器,同时实现最佳的瞬态频率控制,并找到最经济的电力调度频率恢复。为此,我们将强化学习(RL)集成到经典DAI中,从而实现RL-DAI控制。具体而言,我们使用RL学习基于神经网络的控制策略映射从DAI的积分变量的可控功率注入,提供最佳的瞬态频率控制,而DAI固有地确保频率恢复和最佳的经济调度。与现有的方法相比,我们提供了可证明的保证学习控制器的稳定性和扩展的一组允许的成本函数到一个更大的类。在39节点新英格兰系统上的仿真验证了我们的结果。
Frequency control plays a pivotal role in reliable power system operations. It is conventionally performed in a hierarchical way that first rapidly stabilizes the frequency deviations and then slowly recovers the nominal frequency. However, as the generation mix shifts from synchronous generators to renewable resources, power systems experience larger and faster frequency fluctuations due to the loss of inertia, which adversely impacts the frequency stability. This has motivated active research in algorithms that jointly address frequency degradation and economic efficiency in a fast timescale, among which the distributed averaging-based integral (DAI) control is a notable one that sets controllable power injections directly proportional to the integrals of frequency deviation and economic inefficiency signals. Nevertheless, DAI does not typically consider the transient performance of the system following power disturbances and has been restricted to quadratic operational cost functions. This paper aims to leverage nonlinear optimal controllers to simultaneously achieve optimal transient frequency control and find the most economic power dispatch for frequency restoration. To this end, we integrate reinforcement learning (RL) to the classic DAI, which results in RL-DAI control. Specifically, we use RL to learn a neural network-based control policy mapping from the integral variables of DAI to the controllable power injections which provides optimal transient frequency control, while DAI inherently ensures the frequency restoration and optimal economic dispatch. Compared to existing methods, we provide provable guarantees on the stability of the learned controllers and extend the set of allowable cost functions to a much larger class. Simulations on the 39-bus New England system illustrate our results.