An integrated critic-actor neural network for reinforcement learning with application of DERs control in grid frequency regulation

An integrated critic-actor neural network for reinforcement learning with application of DERs control in grid frequency regulation
复制标题

用于强化学习的集成批评者-行动者神经网络以及 DER 控制在电网频率调节中的应用

DOI:
10.1016/j.ijepes.2019.04.011
复制
发表时间:
2019
影响因子:
5.2
通讯作者:
Hu Yu Hen
Hu Yu Hen
中科院分区:
工程技术2区
文献类型:
--
作者:
Sun Jian;Zhu Zhiqin;Li Huaqing;Chai Yi;Qi Guanqiu;Wang Huiwei;Hu Yu Hen

文献摘要

被引文献

相似文献

随着分布式能源在电网中的迅速发展,电力需求的满足和频率调节是控制领域面临的两个主要挑战。然而,这是很难建模的分析,以及设计一个稳定的和最优的控制方案的大规模电网。在DERs的支持下,本文提出了一种行动-批评神经网络,它集成了分布式强化学习控制方案,以补偿电网的频率调节。通过确定性学习算法获得期望控制输出的近似值,提高了系统的短期性能和稳定性。同时,一个长期的战略效用函数估计的综合行动者-评论家神经网络。通过神经网络识别系统状态和控制输出到策略效用函数值的映射,并利用次优控制学习进一步改善系统的长期性能。理论分析保证了稳定性。频率偏差、联络线潮流和长期费用符合一致极限有界性。此外,还计算了长期系统费用的上限。通过两个算例验证了该方法的有效性和优越性。仿真结果表明,在一定条件下,该控制方案比传统的ACTOR-CRITICAL网络控制方案在电网频率调节中具有更好的性能。
As the electronically-interfaced distributed energy resources (DERs) grow rapidly in power grid, power demand satisfaction and frequency regulation are two main challenges in control area. However, it is difficult to model the analysis of a large-scale grid as well as design a stable and optimal control scheme. With the support of DERs, this paper proposes an actor-critic neural network that integrates a distributed reinforcement learning control scheme to compensate frequency regulation of power grid. The short-term performance and stability is improved by a deterministic learning algorithm that is used to obtain the approximation of desired control output. Meanwhile, a long-term strategic utility function is estimated by the integrated actor-critic neural network. The mapping from system state and control output to the strategic utility function value is identified by neural network, as well as utilized in sub-optimal control learning for further improvement of long-term system performance. Theoretical analysis guarantees the stability. Frequency deviation, tie-line power flow, and long-term cost are coincident with uniform ultimate boundness (UUB). In addition, the upper bound of long-term system cost is also reckoned. The effectiveness and advantages of proposed scheme are illustrated in two case studies. The simulation results indicate that the proposed scheme has better performance under certain condition, compared with some actor-critic network control schemes in frequency regulation of power grid.