AdapSafe: Adaptive and Safe-Certified Deep Reinforcement Learning-Based Frequency Control for Carbon-Neutral Power Systems

AdapSafe: Adaptive and Safe-Certified Deep Reinforcement Learning-Based Frequency Control for Carbon-Neutral Power Systems
复制标题

DOI:
10.1609/aaai.v37i4.25660
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Xu Wan;Mingyang Sun;Boli Chen;Zhongda Chu;F. Teng
Xu Wan;Mingyang Sun;Boli Chen;Zhongda Chu;F. Teng
中科院分区:
其他
文献类型:
--
作者:
Xu Wan;Mingyang Sun;Boli Chen;Zhongda Chu;F. Teng

文献摘要

相似文献

随着基于逆变器的可再生能源的日益普及,深度强化学习(DRL)已被提出作为实现未来碳中和电力系统实时和自主控制的最有前途的解决方案之一。特别是,基于DRL的频率控制方法已被广泛研究,以克服基于模型的方法的局限性,如大规模系统的计算成本和可扩展性。然而,基于DRL的频率控制方法的实际实施面临以下基本挑战:1)在学习和决策过程中的安全保证; 2)对动态系统操作条件的适应性。为此,这是第一项提出用于频率控制的自适应和安全认证DRL(AdapSafe)算法的工作,以同时解决上述挑战。特别是,一种新的自校正控制障碍函数的设计,以主动补偿不安全的频率控制策略下的变化的安全约束,从而实现有保证的安全。此外,元强化学习的概念被集成,以显着提高其在非平稳电力系统环境中的适应性,而不牺牲安全成本。基于GB 2030电力系统进行了实验,结果表明,所提出的AdapSafe在训练和测试阶段均具有上级的安全性,并对系统参数的动态变化具有较强的适应性。
With the increasing penetration of inverter-based renewable energy resources, deep reinforcement learning (DRL) has been proposed as one of the most promising solutions to realize real-time and autonomous control for future carbon-neutral power systems. In particular, DRL-based frequency control approaches have been extensively investigated to overcome the limitations of model-based approaches, such as the computational cost and scalability for large-scale systems. Nevertheless, the real-world implementation of DRLbased frequency control methods is facing the following fundamental challenges: 1) safety guarantee during the learning and decision-making processes; 2) adaptability against the dynamic system operating conditions. To this end, this is the first work that proposes an Adaptive and Safe-Certified DRL (AdapSafe) algorithm for frequency control to simultaneously address the aforementioned challenges. In particular, a novel self-tuning control barrier function is designed to actively compensate the unsafe frequency control strategies under variational safety constraints and thus achieve guaranteed safety. Furthermore, the concept of meta-reinforcement learning is integrated to significantly enhance its adaptiveness in non-stationary power system environments without sacrificing the safety cost. Experiments are conducted based on GB 2030 power system, and the results demonstrate that the proposed AdapSafe exhibits superior performance in terms of its guaranteed safety in both training and test phases, as well as its considerable adaptability against the dynamics changes of system parameters.