Multi-Agent Safe Policy Learning for Power Management of Networked Microgrids

Multi-Agent Safe Policy Learning for Power Management of Networked Microgrids
复制标题

DOI:
10.1109/tsg.2020.3034827
复制
发表时间:
2019-07
影响因子:
9.6
通讯作者:
Qianzhi Zhang;K. Dehghanpour;Zhaoyu Wang;F. Qiu;Dongbo Zhao
Qianzhi Zhang;K. Dehghanpour;Zhaoyu Wang;F. Qiu;Dongbo Zhao
中科院分区:
工程技术1区
文献类型:
--
作者:
Qianzhi Zhang;K. Dehghanpour;Zhaoyu Wang;F. Qiu;Dongbo Zhao

文献摘要

被引文献

相似文献

提出了一种基于监督的多智能体安全策略学习方法(SMAS-PL),用于配电网微网的优化电源管理。虽然无约束强化学习(RL)算法是可能无法满足电网运行约束的黑箱决策模型,但我们提出的方法考虑了交流潮流方程和其他运行限制。因此,训练过程利用操作约束的梯度信息来确保最优控制策略函数产生安全可行的决策。此外,我们开发了一种基于共识的分布式优化方法来训练代理的策略功能,同时维护MG的隐私和数据所有权边界。经过训练,学习到的最优策略函数可以被MG安全地用于调度他们的本地资源,而不需要从头开始解决复杂的优化问题。通过数值实验验证了该方法的有效性。
This article presents a supervised multi-agent safe policy learning (SMAS-PL) method for optimal power management of networked microgrids (MGs) in distribution systems. While unconstrained reinforcement learning (RL) algorithms are black-box decision models that could fail to satisfy grid operational constraints, our proposed method considers AC power flow equations and other operational limits. Accordingly, the training process employs the gradient information of operational constraints to ensure that the optimal control policy functions generate safe and feasible decisions. Furthermore, we have developed a distributed consensus-based optimization approach to train the agents’ policy functions while maintaining MGs’ privacy and data ownership boundaries. After training, the learned optimal policy functions can be safely used by the MGs to dispatch their local resources, without the need to solve a complex optimization problem from scratch. Numerical experiments have been devised to verify the performance of the proposed method.