Reinforcement Learning for Energy-Storage Systems in Grid-Connected Microgrids: An Investigation of Online vs. Offline Implementation

Reinforcement Learning for Energy-Storage Systems in Grid-Connected Microgrids: An Investigation of Online vs. Offline Implementation
复制标题

DOI:
10.3390/en14185688
复制
发表时间:
2021-09
期刊:
影响因子:
3.2
通讯作者:
Khawaja Haider Ali;M. Sigalo;Saptarshi Das;E. Anderlini;A. Tahir;M. Abusara
Khawaja Haider Ali;M. Sigalo;Saptarshi Das;E. Anderlini;A. Tahir;M. Abusara
中科院分区:
工程技术4区
文献类型:
--
作者:
Khawaja Haider Ali;M. Sigalo;Saptarshi Das;E. Anderlini;A. Tahir;M. Abusara

文献摘要

相似文献

由可再生能源、电池存储和负载组成的并网微电网需要一个适当的能源管理系统来控制电池的运行。传统上,使用24小时的负荷需求和可再生能源(RES)产生的预测数据,使用离线优化技术来优化电池的操作,其中在一天开始之前确定电池动作(充电/放电/空闲)。强化学习(RL)最近被认为是这些传统技术的替代方法,因为它能够使用真实数据在线学习最优策略。在文献中已经提出了两种研究RL的方法,即。离线和在线。在脱机RL中,代理使用预测的发电和负载数据来学习最优策略。一旦实现融合,电池命令就会实时发送。这种方法与传统方法类似,因为它依赖于预测数据。另一方面,在在线RL中,代理通过使用真实数据与系统实时交互来学习最优策略。本文考察了这两种方法的有效性。将不同标准差的高斯白噪声添加到实际数据中,生成合成预测数据,以验证该方法的有效性。在第一种方法中,预测数据由离线RL算法使用。在第二种方法中,在线RL算法与真实的流数据实时交互,并使用真实的数据来训练代理。当两种方法的能量成本进行比较时,发现当实际数据与预测数据之间的差异大于1.6%时,在线RL提供了比离线方法更好的结果。
Grid-connected microgrids consisting of renewable energy sources, battery storage, and load require an appropriate energy management system that controls the battery operation. Traditionally, the operation of the battery is optimised using 24 h of forecasted data of load demand and renewable energy sources (RES) generation using offline optimisation techniques, where the battery actions (charge/discharge/idle) are determined before the start of the day. Reinforcement Learning (RL) has recently been suggested as an alternative to these traditional techniques due to its ability to learn optimal policy online using real data. Two approaches of RL have been suggested in the literature viz. offline and online. In offline RL, the agent learns the optimum policy using predicted generation and load data. Once convergence is achieved, battery commands are dispatched in real time. This method is similar to traditional methods because it relies on forecasted data. In online RL, on the other hand, the agent learns the optimum policy by interacting with the system in real time using real data. This paper investigates the effectiveness of both the approaches. White Gaussian noise with different standard deviations was added to real data to create synthetic predicted data to validate the method. In the first approach, the predicted data were used by an offline RL algorithm. In the second approach, the online RL algorithm interacted with real streaming data in real time, and the agent was trained using real data. When the energy costs of the two approaches were compared, it was found that the online RL provides better results than the offline approach if the difference between real and predicted data is greater than 1.6%.